Beta DeployAngel is in beta. Every feature is free while it lasts, with no credit card. What beta means →

DeployAngel Beta

Blog

Your health check isn't evidence your deploy worked

Jordan Owens ·

I'm building DeployAngel, which watches each Rails deploy in production and decides when it's safe to stop watching it. It runs on itself, and on its second day its own releases started clearing at High confidence, 17 minutes after each deploy. That felt good until I looked at what the clearances rested on.

The clearance report said "2 of 2 routes ran." One was the agent's telemetry endpoint. The other was GET /up.

What I found

DeployAngel runs on DigitalOcean App Platform, which checks /up about every 10 seconds to decide whether to send the container traffic. Over one day that came to 8,689 health-check requests, against a few hundred for the dashboard pages people actually use.

So when a release had been live for 17 minutes and had served 119 requests, most of those were the platform asking "are you up?" and getting "yes." None of the pages a person uses had to run for the release to clear at High confidence. If I'd shipped a broken dashboard, the health check would have kept answering 200, and the numbers would have looked fine.

What /up actually tells you

Every Rails app generated since 7.1 has this route:

get "up" => "rails/health#show", as: :rails_health_check

Rails::HealthController returns 200 if the app booted without raising, and 500 if it didn't. It doesn't touch the database, Redis, your job queue, or any page you've written. That's exactly what a load balancer needs: is this process ready to take traffic? It's not an answer to "did this release work?"

The trouble is that it's often the busiest route in a small app, and it's always fast and always successful. Anything that measures your app as a whole counts it unless you tell it not to.

Three ways it hides a bad release

It meets your minimums. Any check that waits for "enough traffic" before judging a release gets there quickly when a load balancer is generating six requests a minute per instance. The release looks well exercised when nothing a user touches has run.

It dilutes your error rate. Say a release breaks one page and 10 of your 1,000 real requests in the next hour fail. That's a 1% error rate, enough to trip most alerts. Add 9,000 health checks in the same hour and the app-wide rate is 0.1%, which usually trips nothing.

It drags your p95 down. If 90% of your requests are 2 ms health checks, your app-wide p95 isn't the 95th percentile of anything a user sees. The health checks fill the fastest 90%, so the top 5% of all requests is the slower half of your real ones: the "p95" on your dashboard is your real pages' median. A release that doubles the time of your slowest pages barely moves it.

There's a fourth, smaller version of the same problem: bots. Over the last day DeployAngel got 175 requests that matched no route, for things like /wp-admin and .env. They're answered with a fast 404 and they pad the request count and pull latency down the same way.

What to do about it, whatever you use

What changed in DeployAngel

The agent no longer records health checks at all. It recognizes Rails' /up wherever it's mounted, the OkComputer, health_check, and rails-healthcheck gems, and a lambda or Rack app at a conventional path such as /healthz. A health check served by your own controller can be listed:

DeployAngel.configure do |config|
  config.ignored_routes = [ "GET /healthz" ]
end

Requests no route matched are still recorded, but an unrouted 4xx no longer counts toward the app's request count, error rate, or latency. An unrouted 5xx still counts, since something broke. The server also drops /up from older agents, and I rewrote the telemetry already stored, so baselines and new releases are measured the same way.

Here's the same app, the same 17 minutes after a deploy, before and after:

Before the fixAfter the fix
Routes countedBefore: 2 of 2: /up and the telemetry APIAfter: 1 of 1: the telemetry API
RequestsBefore: 119After: 49
Job runsBefore: 186After: 184
ConfidenceBefore: HighAfter: Medium

Medium is the right answer. In 17 minutes on a quiet dashboard nobody signed in, so the release is cleared on the API the agents call, the jobs that ran, and the scheduled jobs that ran on time, and the report says so. Dashboard pages that haven't run yet are checked the first time someone uses them. High confidence there was a health check talking.

The general lesson I keep relearning while building this: a verdict is only as good as what it's counting, so it should always say what it counted.

About DeployAngel

DeployAngel verifies each Rails deploy in production. A small gem reports aggregated telemetry once a minute, and each release is compared with the last healthy one and given a verdict: cleared, failed, or not cleared yet, with what it rests on. CI and coding agents can wait for that verdict through a CLI or an MCP server. It's free during the beta.