Klaxo watches your metrics, your logs, your cluster — and your support queue. When something breaks it raises the incident itself, and the answer is already written.
checkout-api · http_5xx_ratio · 02:30–03:58 UTCReal outage. Raised as INC-3312, sev-2. checkout-api 5xx on POST /v2/capture went from 0.31% to 7.4% in the 17 minutes after the 03:24 config push to pricing-svc. Confined to the payment-capture path — browse and cart are nominal.
Cause is upstream, not checkout-api itself: push 8f2c1d took the pricing-svc pool from 200 connections down to 40, the pool now saturates at 40/40, and checkout-api fails closed instead of falling back to the cached quote. 11 of 14 checkout-api pods report PoolTimeout; the three still holding a warm connection are clean.
Corroborated by the support queue — 3 tickets in 9 minutes from distinct users, all on the same payment step: SUP-4483 opened it at 03:33, then SUP-4485 at 03:37, then SUP-4488 at 03:41 — the ticket linked to this incident. Breadth confirmed, so this was auto-raised rather than held.
One ticket, on a night with twenty other one-offs. Nothing about it looks like an outage yet.
Different words. Same shape.
The open queue is breadth-checked for tickets that agree — then the claim goes to a pass whose job is to refute it.
No human nudge. Nobody was awake to give one.
Twenty-four minutes after the first customer did.
Ordinary tickets look ordinary. Klaxo breadth-checks the open queue for the ones that agree, and only a claim that survives a deliberately skeptical second pass is allowed to raise anything. An unconfirmed one quietly becomes a single engineering ticket instead of waking someone.
checkout-api 5xx at 7.4% vs 0.31% baseline
A week of the entity's own behaviour. Not a number somebody guessed a year ago.
Nothing to write. Nothing to tune. The metric is compared against itself.
Four times its own baseline, and still climbing.
A noisier lane needs a stricter gate — so a fire is checked against the queue and challenged by a second pass before it can raise anything.
Coupon service returning 500s on ~8% of carts since 02:58, right after a config push to pricing-svc. Real, customer-facing, not self-healing.
Metric, log and cluster-state rules on a loop, deduped on the firing fingerprint so a restart never doubles an incident.
Describe it in plain language, then see whether it would have caught the outages you remember — before it can page anyone.
No authored number at all. Machine-seeded sweeps land as proposals for a human to arm.
Every ticket resolves to exactly one of four outcomes: reply to the reporter, file an engineering ticket, raise an incident, or name the manual step a human has to take. Drafted, evidenced, waiting on one click.
pricing-svc returns $0 tax on 3 postal codes
eng ticket
held
Hi Dana — your export EXP-71204 was stuck behind an
export-worker backlog, not anything wrong on your side. That backlog is
draining now and yours is second in the queue.
Expect it by 04:10. I’ll update this ticket the moment it finishes — nothing is needed from you.
checkout-api 5xx 7.4% vs 0.31% baseline
incident
raised
A tool-using agent queries logs and metrics, walks live cluster state, reads the recent commits and the runbooks, and returns a verdict — cause, blast radius, whether it is customer-facing. Nobody writes a prompt.
Rule R-014: checkout-api 5xx above 5% for three consecutive 30s windows.
5xx ratio 0.31% → 7.4% across 12 consecutive windows — 24× the same hour last week.
asyncpg.PoolTimeoutError 2,847 times in 9 minutes, on 11 of 14 pods.
deploy/checkout-api unchanged — 14/14 pods ready, zero restarts in the window.
3 tickets on the same payment step, 03:33 → 03:41: SUP-4483, SUP-4485, SUP-4488.
pricing-svc pool cap 200 → 40 in 8f2c1d, pushed 03:24 — the one change correlated with the window.
Same signature as INC-2980 on 19 Jun, which was fixed by reverting the push.
The 03:24 push cut pricing-svc’s pool cap from 200 to 40. It saturates at 40/40, so checkout-api blocks waiting on a quote and returns 5xx. Not traffic: RPS flat at 1.02x. Not a checkout rollout: zero restarts, no image change in 6h.
Authored rules on one side, a baseline that needs no threshold on the other. Plus the support queue, which no monitoring tool reads.
Corroborated across the queue, then handed to a pass whose job is to refute it. Only survivors get raised.
Logs, metrics, live cluster state, the recent commits, the runbooks. Nobody writes a prompt.
Raise, reply, file, review, propose a fix. How far it goes alone is a dial you set.
Every outcome becomes a fact the next investigation already knows. That is why the second one is faster.
Four settings. Start at zero, watch what it would have done, move the dial when it has earned it. One deliberate exception: a verified outage is always raised, even when everything else waits for a click.
Similar incidents collapse into one cause, so four tickets read as one problem — and the gaps get attributed. Your retro brief is written before the meeting starts.
18 of the 47 are one signature — catalog 5xx spike, every one between 02:00 and
04:30, every one downstream of the pricing-svc price sync. RCA graded strong; the fix
is still unshipped. Four signatures cover 40 of the 47; the other 7 never recurred.
These 47 are real — one production instance, 30 days, every incident Klaxo raised or was handed. The outage walked through earlier on this page is a composite drawn from them.
No migration. No re-platforming.
A promised turnaround is a commitment somebody has to keep. So it comes out of a table, checked against the replies your team has already sent. When the request doesn't match cleanly, the reply carries no date at all — silence beats promising two days on something that takes ten.
And when they do, the investigation should already be finished.