What an SLA actually promises

May 25, 2026

I’ve inherited more than one service level agreement that had been green for two years straight while the business complained constantly.

That combination should be alarming, and usually isn’t, because everyone involved has quietly agreed not to look at it. The report says 99.9%. The users say it’s down all the time. Both are true, which means the measurement is measuring something nobody cares about.

Availability of what, exactly

Most bad SLAs fail on the definition, not the number.

If you measure whether the server responds to a ping, you’ll get a beautiful number that has nothing to do with whether anyone could do their job. If you measure whether a user in a branch office could complete a transaction, you’ll get an uglier number that’s actually worth reporting.

The questions I ask of any service level:

  • Measured from where? From inside the data centre, or from where the user sits. These are very different services.
  • Measured against what? A transaction the business recognises, or a component the business has never heard of.
  • Whose clock? Does the incident start when monitoring alerted, when the service desk logged it, or when the first user noticed? Most reporting quietly uses the second one, which excludes the worst part of the outage.
  • What’s excluded? Planned maintenance almost always is. If your maintenance window is Saturday night and your users are global, you’ve excluded a real outage for a real population.

None of this is about being pessimistic. It’s about the number meaning the same thing to the person reporting it and the person reading it.

Availability is the least interesting number

The service levels that changed behaviour for me were rarely availability.

On banking applications supporting several billion in managed assets, the number that mattered wasn’t uptime — it was recovery. How long from detection to service restored. Uptime is an outcome you mostly can’t control in the moment. Recovery time is a thing you can practise, staff, and improve. Redesigning the incident recovery process moved SLA performance about 20%, and none of that came from making the systems more reliable.

Worth reporting alongside availability:

  • Time to detect. How long it was broken before anyone knew.
  • Time to restore. Detection to service back.
  • Repeat rate. How much of this month you’d seen before.

Report the misses

The strongest thing you can do for an SLA’s credibility is show a breach, in the same report, in the same format, without a paragraph of explanation attached.

An SLA report that’s green every month gets skimmed. One that occasionally goes red — with the cause and the fix next to it — gets read, because the reader learns it tells the truth. That credibility is the whole asset. It’s what lets you go to the same audience later and ask for money.

Tell us what is breaking, what is slow, or what you are afraid to touch.

Every engagement starts with a conversation about outcomes, not hours. If we are not the right fit, we will say so.