Skip to content
Case Study

The Renewal Gets Signed Because the Cutover Is Scary

Atin Agarwal · Aug 21, 2026 · 9 min read
The Renewal Gets Signed Because the Cutover Is Scary featured image
migrationvendor-lock-inalert-rulesparallel-runcase-study
Share:

The renewal quote arrives six weeks before the date. Everyone in the room agrees it is too much money for what it does. Everyone agreed last year, and the year before.

Then somebody asks what moving would involve, and the answer is: rebuild four hundred alert rules, forty dashboards and eleven integrations, none of which are documented, while the team is already committed for the quarter — and accept a window in the middle where the old system is being decommissioned and the new one is not yet trusted.

The renewal gets signed. It is not weakness or inertia. Given how the choice was framed, it is the correct decision, and it will be the correct decision again next year, for the same reason, forever.

The Lock-In Is Not the Contract

Nothing in the agreement prevents leaving. What prevents leaving is everything that accumulated on top of it.

Configuration nobody can account for. Hundreds of alert rules written over four years by people who have mostly moved on. Some are disabled. Several are duplicates with slightly different thresholds. A meaningful number reference metrics that are no longer emitted and cannot fire at all. Nobody can tell you which is which, so the migration gets priced as if all four hundred must be faithfully reproduced.

Knowledge that exists in no file. Which alerts matter. Which ones everybody ignores. Which threshold was set to a strange number because of an incident in 2023 that nobody has written down. That knowledge is the actual asset, it lives in two people’s heads, and rebuilding it is the part of the migration that cannot be estimated.

A gap that is genuinely dangerous. This is the honest core of the fear. In a straight cutover there is a period where the old system is winding down and the new one is untuned — thresholds not yet calibrated, routing not yet proven, nobody sure which alerts to trust. If an incident lands in that window and is missed, it becomes the story of the migration, and everyone involved knows it.

And the switching cost is paid by engineering while the saving lands in finance. The team doing the work absorbs a quarter of undifferentiated effort; the benefit appears on somebody else’s line. That asymmetry decides more renewals than any feature comparison, and it is the same one behind the $700K infrastructure illusion: the spend is visible, and the reason it persists is organisational rather than technical.

Two timelines. The upper timeline, labelled cutover, shows a green bar for the old tool that ends, followed by a red hatched gap where neither system is trusted, followed by a bar for the new tool that starts untuned. The lower timeline, labelled parallel run, shows both bars overlapping for several weeks with no gap between them, and a marker where the old tool is switched off after the new one has been verified against it. A strip along the bottom notes that the risk of a migration is the gap, and a parallel run has no gap.
The thing everyone is afraid of is the red section. It is an artefact of cutting over, not of migrating.

What Turns Up When You Actually Look

We have run this migration many times, and the first pass over the outgoing system produces the same findings.

  • The rule set is largely dead. Duplicates, disabled rules, rules on decommissioned services, rules that cannot fire. A third to a half of the list routinely does not survive first contact with the question “what would this alert someone about, and would they act?”
  • The dashboards are mostly unopened. Access logs settle this quickly. A handful get used during incidents; the rest were built for a project that ended. Migrating all of them is the single largest line in most estimates and the least justified.
  • There is no definition of done. Without an inventory of what is actually running, “we have migrated the monitoring” cannot be verified — only asserted. You cannot port coverage you cannot enumerate, which is why an accurate inventory turns out to be the first task in a tooling migration as well as the first task in everything else.
  • The integrations are underestimated. Ticketing, chat, paging, the status page, the SSO configuration, the one webhook feeding a spreadsheet somebody in support depends on. Individually trivial, collectively the part that overruns.
  • Nobody knows what the tool is actually being paid for. Host counts that predate an autoscaling change, retention set at a tier nobody chose, custom metrics ingested by a service that was retired. The bill is rarely reconciled against usage, for the same reasons covered in getting observability costs under control.
  • The estimate assumes reproduction rather than replacement. Every plan we are shown starts from the old tool’s configuration list. Starting there guarantees you rebuild the accumulated mess in a new place, at full cost, and inherit it for another four years.

So the quoted migration cost is usually wrong by a wide margin, in both directions. Reproducing everything costs more than anyone budgets. Migrating what is actually load-bearing costs considerably less.

Run Both. Then There Is No Gap.

The fear is specific and it has a specific answer. The risk in a migration is not the new tool being bad; it is the interval where neither is trusted. Remove the interval and the argument collapses.

Instrument the same estate twice, and leave the old one running. The new system collects from the same infrastructure, in parallel, with its alerts routed to a shadow destination that pages nobody. Cost overlaps for a period. That overlap is the entire price of eliminating the risk, and it is far cheaper than another year of the renewal.

Then compare, with evidence. Every time the old system fires, did the new one? Every time the new one fires, was it real? Run that for a full cycle — long enough to include a deploy, a traffic peak and at least one genuine incident. At the end you are not asserting that coverage is equivalent; you have a record showing it.

Migrate coverage, not configuration. Start from the inventory of what exists and what it does for customers, and write the rules that protect it. Do not start from the old rule list. The old rule list is a history of everything anyone ever worried about, including the things that turned out not to matter, and it is the mechanism by which four years of accumulated noise gets carried into a fresh system.

Cut over per service, not all at once. Each service moves when its shadow alerts have been correct for a fortnight. A rollback is one service, not a project.

Decommission last. The old system stays until the new one has carried a real incident end to end. That is the only test that counts, and it is cheap to wait for.

CutoverParallel run
Coverage gapWeeks, in the middleNone
Starts fromThe old rule listThe inventory
Dead rulesFaithfully reproducedLeft behind
Equivalence isAssertedEvidenced, alert by alert
Rollback unitThe whole migrationOne service
Old system off whenThe new one is deployedIt has carried an incident
CostOne quarter of engineeringA few weeks of double licence

The Part a Tool Can’t Do

Every vendor will import your dashboards. Some will convert alert rules automatically, and the conversion mostly works. None of that touches the actual problem, because the actual problem is deciding which of the four hundred rules deserves to exist in the new system.

That is a judgement per rule: what does this protect, would anyone act on it, and is the threshold still right for the traffic you have now. It requires knowing the architecture and it requires being willing to delete things — which is uncomfortable, because deleting an alert feels like removing a safeguard even when the alert has been firing into a muted channel for two years. Nobody makes that call confidently without having run the estate.

An automated import is worse than nothing here, in a specific way: it makes the migration look finished while carrying the entire accumulated mess across intact. A year later the new tool has the same four hundred rules, the same muted channel, and the same conversation at renewal — alert fatigue restored to factory settings.

And somebody has to own the parallel period. Watching two systems and adjudicating the differences is real work for several weeks, done by people who understand what should have fired. That is the labour the migration actually consists of, and it is the reason the project keeps not happening — not the tooling, which is the easy half.

IOanyT Innovations has been running client infrastructure for over ten years, across more than 150 engagements, and most of those estates came to us on somebody else’s tooling with a renewal date approaching.

Vigil runs in parallel with whatever you have now, on Prometheus and Grafana inside your own cloud accounts. We start from discovery rather than from your old rule list, shadow your existing alerts until the comparison is boring, migrate service by service, and hold the on-call throughout — including the overlap. You keep the incumbent until you have watched us take a real incident.

The gap everyone is afraid of is ours to carry, not yours.

See what’s included in monitoring →

Start with a free infrastructure assessment →

Atin Agarwal

About the Author

Atin Agarwal

Founder, IOanyT

Atin has spent 25+ years building and operating infrastructure systems across 150+ client engagements. He writes about the gap between what monitoring tools promise and what actually keeps systems healthy.

See outcome ownership in action

Your infrastructure deserves more than a dashboard. Schedule a demo to see how Vigil handles the monitoring — and the 2 AM pages.