Beta DeployAngel is in beta. Every feature is free while it lasts, with no credit card. What beta means →

DeployAngel Beta

Blog

Verification debt: shipping faster than you can check

Jordan Owens ·

Over three days this week I deployed the same app 43 times. Most of that code was written by a coding agent, with me reviewing it. Once the app was past its first day, the median time between my deploys was 14 minutes.

I know that number because the app is a tool that decides whether each deploy worked, and the fastest it has ever been able to say yes is 17 minutes. Of the 18 releases after that first day, 9 were replaced by the next one before the check on them had finished.

I don't think that's unusual any more. I think it's where a lot of us are heading, and it has a cost that doesn't show up anywhere until it does. I've started calling it verification debt.

Writing code got cheap. Checking it didn't.

For most of my career the slow part of shipping was writing the code. Review, CI, and a look at the dashboards after a deploy were sized for a few releases a day, made by people who had spent hours with the change and had a feel for what might break.

Coding agents moved the bottleneck. A change that took an afternoon takes twenty minutes, and the agent's natural stopping point is "deployed." But nothing about production got faster. Real users still arrive at the rate they arrive. The nightly job still runs at night. A slow page still needs enough visits before its p95 means anything.

So every release that nobody checked becomes a small debt: a claim that it works, made on nobody's evidence. It builds up silently, because from the outside "deployed" and "done" look exactly the same.

What the debt looks like

The first time it cost me something was a slowdown. On the first day a release failed its check: p95 latency had doubled, from 1,102 ms to 2,148 ms. The natural move was to look at that release's diff.

But the slowdown didn't start with that release. The three releases before it were already that slow. Each had been replaced by the next deploy before there was enough traffic to judge it, so nothing had ever said a word about them. The failure landed on whichever release happened to be the first one checked, and the search for the cause had to start four releases back.

That's what verification debt does when it comes due. You don't just find out late. You find out about the wrong release, and you pay interest searching further back.

Why the usual signals don't settle it

Each of the things we normally lean on answers a smaller question than it seems to.

None of these are useless. They're just answers to "did it deploy?" and "did anything explode?", and the question that matters is "is it safe to stop watching?"

When is it safe to stop watching a deploy?

This is the checklist I've ended up with. It works whether the one checking is you, a script, or a tool.

  1. Compare with the previous good release, not with "normal." Last Tuesday's traffic isn't today's. The release before this one, on similar traffic, is the fairest baseline you have.
  2. Compare route by route and job by job. App-wide numbers average a broken page away. A 1% error rate on checkout can be a 0.1% error rate on the app.
  3. Count only real evidence. Leave out health checks, bot probes, and anything else that's always fast and always succeeds.
  4. Wait for what's supposed to happen. If a scheduled job should run in the next hour, the release isn't done until it has. Same for the business events you count on, like sign-ups or orders.
  5. Name what hasn't run yet. "Nothing failed" and "we saw the password reset work" are different statements. Write down which parts of the app the verdict actually covers.
  6. Scale the wait to your traffic. A busy app can tell you in fifteen minutes. A quiet one might need hours, or a smoke test that visits the pages that matter.
  7. Don't skip releases. If you deploy again before a verdict, the next verdict has to cover both releases, compared with the last one that was actually good. Otherwise the debt just disappears from view.
  8. Let AI explain, not decide. Use it to read a failure and point at a cause. Keep the pass or fail on rules you can read, so it means the same thing every time.

Paying it down

The fix isn't to deploy less. Shipping small changes often is still the right instinct; small releases are easier to judge and easier to roll back. The fix is to put a verdict between "deployed" and "done" for every release, and to write down what that verdict rested on.

For a team, that means a release isn't finished when it's live. It's finished when something has compared it with the last good release, on real traffic, and said so. For an agent, it means the task isn't done when the deploy command exits. The agent waits for the verdict, and if the verdict is "not yet," it says that instead of "done."

And when you do deploy faster than verdicts arrive, which you will, track it. The number of releases since the last one that was actually verified is a better measure of risk than the number of deploys.

About DeployAngel

I'm building DeployAngel to do this checklist automatically for Rails apps. A small gem reports aggregated telemetry once a minute, and each release is compared with the last cleared one and given a verdict, with what it rests on and how many releases are waiting on it. CI and coding agents can wait for that verdict through a CLI or an MCP server. It's free during the beta.