USCodeHub All articles
Career & Salaries

Always On, Always Exhausted: The On-Call Culture That's Costing You Your Best People

USCodeHub
Always On, Always Exhausted: The On-Call Culture That's Costing You Your Best People

The pager goes off at 3:17 a.m. Your senior engineer — the one who knows where every skeleton is buried in your codebase, the one who's been on the team for four years — rolls over, grabs their phone, and starts triaging an incident they'll spend the next two hours debugging. By morning, they're back in meetings, expected to perform at full capacity.

This happens again Thursday. And again the following Monday. And somewhere around month eight of a rotation that was supposed to be shared but somehow always seems to land hardest on the people with the most context, something shifts. They start looking at LinkedIn. They update their resume. And then they're gone, taking four years of institutional knowledge with them.

On-call culture is one of the most normalized forms of engineer burnout in the US tech industry. And it's hiding in plain sight.

The Sleep Debt Is Real and It Compounds

The research on sleep deprivation and cognitive performance is not ambiguous. A person operating on fragmented sleep performs meaningfully worse on complex problem-solving tasks — exactly the kind of tasks software engineers are paid to do. Debugging a distributed system failure requires working memory, pattern recognition, and the ability to hold multiple hypotheses simultaneously. All of those degrade significantly when you're running on four hours of interrupted sleep.

Cumulative sleep debt doesn't reset with a single good night. Engineers who are regularly paged during overnight hours accumulate a cognitive deficit that affects their daytime work in ways that are hard to measure but very real. They make more mistakes. They move more slowly. They have less patience for the kind of deep work that produces good code.

Most companies track engineer velocity through sprint completion and story points. Almost none of them track the performance impact of their on-call schedule. The cost is invisible in the metrics, but it's absolutely showing up in the output.

Senior Engineers Bear the Weight Nobody Talks About

On-call rotations are theoretically distributed. In practice, there's a consistent pattern in engineering organizations: the engineers with the most context get paged the most, because they're the ones who can actually resolve incidents quickly.

This creates a perverse dynamic. Your best people — the ones who understand the system deeply enough to fix things fast — end up carrying a disproportionate on-call burden. Junior engineers might be on the rotation, but when something serious breaks at 2 a.m., the escalation path almost always leads back to the same three or four people.

Those people know it. They feel it. And when they eventually burn out and start evaluating their options, they find that companies competing for their experience are often offering roles with no on-call requirements at all. The market has responded to this problem even if individual engineering organizations haven't.

The Context-Switching Tax Is Bigger Than You Think

Beyond sleep, there's the cognitive cost of interruption. An engineer who gets paged during the workday — or who spends part of their evening handling a minor incident — doesn't just lose the time the incident takes. They lose the deep work session they were in, the mental state they'd built up around a complex problem, and often an hour or more of recovery time before they can return to focused work.

Paul Graham wrote about maker schedules versus manager schedules more than a decade ago. The insight still holds. Engineers need long, uninterrupted blocks of time to do their best work. On-call duty is structurally incompatible with that, because the entire point of being on call is to be interruptible at any moment.

When you put your best engineers on a rotation that keeps them in a state of low-grade alertness, you're not just affecting their nights — you're degrading the quality of everything they produce during the day.

What Actually Works: Rethinking the Model

Some engineering organizations have moved toward dedicated incident response teams — specialists whose primary role is triage and initial response, freeing domain engineers from the interruption cycle. This model is more common in larger organizations, but it's worth considering even at mid-size scale. A rotation of four to six people focused entirely on incident response can dramatically reduce the burden on the broader engineering team.

Better alerting is the lower-hanging fruit. Most on-call fatigue is driven not by genuine production emergencies but by noisy, poorly calibrated alerts that page engineers for conditions that resolve themselves or don't actually require immediate human intervention. An alert audit — reviewing every page from the last 90 days and classifying it as actionable, noisy, or unnecessary — typically reveals that a significant percentage of overnight pages were for issues that could have waited until morning or been handled automatically.

Alert severity tiers, clear escalation paths, and ruthless pruning of low-signal noise can cut overnight pages dramatically without any change to the underlying system reliability.

Psychological safety matters too. Engineers who feel they can flag on-call burnout without it being perceived as weakness or lack of commitment are more likely to raise the problem before they're already halfway out the door. Regular retrospectives on on-call burden — separate from incident postmortems — create space for that conversation.

The Retention Math Is Simple

Replacing a senior engineer in the US market costs somewhere between 50% and 200% of their annual salary when you account for recruiting fees, lost productivity during the search, onboarding time, and the knowledge that walked out the door with them. On-call burnout is one of the most frequently cited reasons senior engineers leave roles they otherwise liked.

The math is not complicated. Investing in better alerting infrastructure, restructuring rotations to be genuinely equitable, and creating dedicated incident response capacity costs a fraction of what you'll spend replacing the engineers who burn out under the current system.

The pager is going to go off again tonight. The question is whether the person on the other end of it is going to be there six months from now.

All Articles

Related Articles

How Your Automation Pipeline Became the Biggest Security Hole in Your Stack

How Your Automation Pipeline Became the Biggest Security Hole in Your Stack

When AI Writes the Code, Who Learns to Think?

When AI Writes the Code, Who Learns to Think?

Your Code Reviews Are Broken — Here's How to Actually Fix Them

Your Code Reviews Are Broken — Here's How to Actually Fix Them