The on-call schedule is not a staffing plan

September 2, 2026

There is a specific kind of false confidence that comes from a fully populated on-call calendar. Every week covered, every shift assigned, escalation paths documented, the whole thing sitting in PagerDuty or ServiceNow looking like a real plan. Leadership sees it, nods, and considers the operational risk managed.

It isn’t managed. The schedule is an availability artifact. It is not a capability assessment.

Here is what the schedule cannot tell you: whether the person on call this Saturday has ever worked a P1 alone. Whether they understand the dependencies between the payments platform and the upstream data feeds well enough to know where to look first. Whether the runbook they will open at 2 a.m. was written for the system as it existed eighteen months ago or the system as it exists now. Whether the person listed as their escalation has been on a performance plan for six weeks and is quietly on the way out.

I have managed operations teams inside large, structured environments — the kind with formal ITSM programs, dedicated NOCs, tiered support models, and more process documentation than anyone reads. And across those environments, the gap between the on-call roster and actual incident-response capability was almost always larger than anyone wanted to admit. Not because people were incompetent. Because the roster was maintained and capability was assumed.

Those are two different activities, and most organizations only do one of them.

Capability has a few components that a schedule simply cannot capture. The first is familiarity with the specific failure modes of the specific systems that person supports. This sounds obvious. In practice, it means an engineer who knows your payment processing pipeline cold might be nearly useless on call for your data warehouse team — and yet both names appear in rotations with identical authority on paper. The schedule flattens that distinction. A staffing plan wouldn’t.

The second component is recency. On-call rotation math often spreads coverage so thin that a given engineer might take a shift once every three or four weeks. That is fine if the environment is stable and the engineer is actively engaged with the system between shifts. It is a problem if they spend those three weeks heads-down on development work and haven’t touched an incident since the last rotation. Muscle memory decays. Command of the current state of a system decays faster than most people track.

The third component is load. A schedule tells you a name is available. It does not tell you that person has been carrying two projects and a backfill gap for sixty days and is operating on a level of fatigue that makes complex diagnostic reasoning genuinely harder. Incident response under cognitive load is not the same as incident response rested. The schedule does not model that. Neither does most operational risk thinking.

The fix is not complicated, but it requires treating on-call readiness as a recurring operational question rather than a one-time scheduling problem. That means periodic tabletop walkthroughs with the actual people in the rotation — not the whole team, just the people who will be on call in the next two cycles — working through a realistic failure scenario for the systems they own. Not a full-scale DR exercise. A focused, low-overhead conversation: what breaks, what do you check first, where does the runbook take you, who do you call if that path fails. Thirty minutes. You learn a lot.

It also means tracking call history at the individual level, not just aggregate ticket volume. Which engineers are actually taking incidents to resolution? Which ones are escalating everything after the first fifteen minutes? That pattern is information. If it shows up consistently, you have a coverage problem that the schedule is currently hiding.

None of this requires a new tool or a new framework. It requires someone with operational accountability to treat the on-call program as a live system that degrades without maintenance — because that is exactly what it is.

A name in a rotation slot is a phone number. A staffing plan is a capability map. Know which one you actually have.

Tell us what is breaking, what is slow, or what you are afraid to touch.

Every engagement starts with a conversation about outcomes, not hours. If we are not the right fit, we will say so.