Skip to content
Case Study

Reliability Never Reaches the Board Until It's an Outage

Atin Agarwal · Aug 12, 2026 · 9 min read
Reliability Never Reaches the Board Until It's an Outage featured image
reportingleadershipslaincident-metricscase-study
Share:

Open the last four board decks. There is revenue, and its trend. Pipeline, and its trend. Burn, runway, headcount, churn, NPS — each with a number, a comparison against last month, and a person who owns it.

There is nothing about whether the product stayed up.

Not because the board decided reliability was unimportant. Because no one produces the number, so there is nothing to put on a slide, so it is not on the agenda, so nobody asks for it. The absence is self-sustaining and completely undramatic.

What it guarantees is that reliability enters the conversation exactly once: after an outage. Which means the subject is only ever discussed in its worst frame — reactively, defensively, with a customer email on the screen and somebody explaining what went wrong. Decisions made in that room are made under duress. They are usually generous, and they are usually reversed within a quarter, because nothing keeps the topic alive between incidents.

Reliability is a product feature argues that it deserves the same standing as anything else you build. This is the mechanical reason it does not get it: features have a number that reaches leadership every month, and reliability has an anecdote that reaches them twice a year.

Why the Artefact Doesn’t Exist

The data is per-tool and does not add up. Uptime for one service from one provider’s console, availability for another from a different one, incident history in a chat thread. None of it shares a definition, so there is no honest way to combine it into a single figure. Anyone who has tried has discovered that the estate produces several numbers and no total — the same arithmetic problem as monitoring that stops at one cloud.

Nobody has agreed what to measure. Uptime of what — the marketing site, the API, the checkout path? Measured from where, inside or outside? Does planned maintenance count? Does a partial failure affecting eight per cent of users count, and as what fraction of an hour? Every one of those is a legitimate question with no default answer, and the conversation to settle them has never been scheduled.

Producing it by hand costs a day a month. Which is not much, and is also exactly the kind of recurring cost that never gets approved because it has no urgent trigger. It is easy to skip in a busy month, and once it is skipped once, the trend line it existed to produce is broken anyway.

And engineering is quietly reluctant. A reliability number is a number that can be used against the team, particularly if the definitions were never agreed and particularly if the first month it is published happens to be a bad one. That hesitancy is rational and it is rarely stated out loud, which makes it hard to resolve.

A board deck agenda drawn as a list of rows. Revenue, pipeline, burn, headcount and churn each carry a green tick, a current figure and a trend arrow. The final row, reliability, is drawn with a dashed outline and is empty, with no figure and no trend. A note beside it reads that the subject appears on the agenda only when an incident puts it there. A strip along the bottom notes that nobody decided to leave it out.
Every other row has an owner, a number and a direction. The last one has an anecdote, twice a year.

What Turns Up When You Actually Look

Ask a team for last quarter’s reliability figures and the conversation follows a predictable route.

  • There is no incident record. Incidents happened, were handled well, and live in chat threads. Nobody can say how many there were, so the most basic question — is this getting better or worse — has no answer at any level of precision.
  • The uptime figure comes from the monitoring tool’s own view of itself. It reports the availability of the checks that were configured, which is a statement about the monitoring rather than about the product, and it is almost always flattering.
  • No severity taxonomy. Without one, an incident count is meaningless: a brief degradation and a four-hour outage both increment the same integer, so the number cannot be compared to anything, including itself last quarter.
  • MTTR quoted with no definition behind it. From when — the first symptom, the first alert, or the first person acknowledging? To when — mitigated, or fully resolved? Those are four different numbers, and reporting one of them as though it were the whole thing produces a figure that cannot be acted on.
  • A snapshot instead of a trend. Where a number does exist, it is this month’s. One data point is not information about direction, and direction is the only thing a board can actually act on.
  • Nothing connecting reliability to money. Cost is reported by finance, incidents by engineering, churn by the commercial team, and no artefact puts them on the same page — so the argument for investment is always made on principle rather than on evidence.
  • No record of what was done about it. Even where incidents are logged, the follow-up actions are not tracked to completion, so the same cause recurs and there is no document that shows it recurring. This is where blameless postmortems stop short: the writing-up happens and the aggregate never gets assembled.

The pattern is that everything needed exists and nothing is assembled. The alerts fired, the incidents were handled, the fixes shipped. What was never produced is the one page that turns all of that into a direction of travel.

One Page, Same Definitions, Every Month

The instinct is to build something sophisticated. The opposite is correct: consistency is worth far more than depth, because the entire value is in the comparison to last month.

Agree the definitions once, in writing. Which services are customer-facing and therefore in scope. Measured externally, from where customers are. Whether maintenance windows are excluded. How partial failures are counted. Three severity levels, defined by customer impact rather than by cause. This conversation takes an afternoon and it is the whole project — everything after it is arithmetic.

Report the same six things every month. Availability per customer-facing service. Incident count by severity. Time to detect and time to resolve, each with the definition printed next to it. What changed in the estate. What is planned. And the exceptions — the things known to be unprotected and not yet fixed.

Include the ugly month. A report that only appears when the numbers are good is a marketing document and everyone senior can tell. The credibility of the artefact comes entirely from it being unchanged in a bad month.

Put cost on the same page. Infrastructure spend next to reliability outcomes is the only view that makes the trade-off legible, and it is the view that converts a reliability conversation into a business one. It is also the honest answer to the $700K infrastructure illusion — the spend was always visible; what it purchased never was.

Write the second half in sentences. The numbers say what happened. A short paragraph on what it means and what is being done about it is the part leadership actually reads, and the part no dashboard can generate.

No reportA monthly page
Reliability reaches leadershipVia an outageOn a schedule
The number isAbsent, or the tool’s ownExternally measured, per service
Incidents areChat threadsCounted, by severity
MTTR isUndefinedDefined next to the figure
The view isA snapshot after the factA trend
Investment is arguedOn principle, post-incidentOn evidence, on a cadence
Cost sitsIn a different reportOn the same page

The Part a Tool Can’t Do

Producing the figures is automatable. External checks give availability per service. An incident tool gives counts and durations. A cost explorer gives spend. Assembling them on a page is a scheduled job.

Agreeing what they mean is not, and it is the entire difficulty. Deciding that a partial failure affecting a subset of customers counts as a full outage for reporting purposes is a policy decision with commercial consequences, and it has to be made by people who can be held to it. So does deciding which services are in scope — a question that quietly determines whether the number flatters you.

Somebody also has to stand behind the figure in a room. That means owning it when it is bad, resisting the pressure to redefine the metric after a poor month, and being able to explain what moved and why. No tool can do that, and a report nobody owns is a slide that gets skipped.

And the last half — this is what it means, this is what we are doing, this is what we need — is a judgement about priorities. It is the same scarce input as everywhere else in this series: the data was never the constraint.

IOanyT Innovations has been running client infrastructure for over ten years, across more than 150 engagements. The clients who invest in reliability before an incident are, almost without exception, the ones who receive a number about it every month. Not because the number is persuasive. Because it keeps the subject on the agenda between the events that would otherwise be the only time it appears.

Vigil produces that page: availability measured externally per customer-facing service, incidents counted against a severity scale agreed with you, detection and resolution times with their definitions attached, spend alongside outcomes, and a written summary from the team that carried the on-call — including the months we would rather not send.

Your board gets a trend line instead of an anecdote.

See what’s included in monitoring →

Start with a free infrastructure assessment →

Atin Agarwal

About the Author

Atin Agarwal

Founder, IOanyT

Atin has spent 25+ years building and operating infrastructure systems across 150+ client engagements. He writes about the gap between what monitoring tools promise and what actually keeps systems healthy.

See outcome ownership in action

Your infrastructure deserves more than a dashboard. Schedule a demo to see how Vigil handles the monitoring — and the 2 AM pages.