Klaxo watches your metrics, your logs, your cluster — and your support queue. When something breaks it raises the incident itself, and the answer is already written.
checkout-api · http_5xx_ratio, 90 min windowReal outage. Raised as INC-3312, sev-2. checkout-api 5xx on POST /v2/capture went from 0.31% to 7.4% in the 17 minutes after the 03:24 config push to pricing-svc. Confined to the payment-capture path — browse and cart are nominal.
Cause is upstream, not checkout-api itself: push 8f2c1d took the pricing-svc pool from 200 connections down to 40, the pool now saturates at 40/40, and checkout-api fails closed instead of falling back to the cached quote. 11 of 14 checkout-api pods report PoolTimeout; the three still holding a warm connection are clean.
Corroborated by the support queue — 3 tickets in 9 minutes from distinct users, all on the same payment step (LSD-4488 leading). Breadth confirmed, so this was auto-raised rather than held.
8f2c1d · verdict written 03:42:34, before first human ack · last sweep 03:00 IST
One ticket, on a night with twenty other one-offs. Nothing about it looks like an outage yet.
Different words. Same shape.
The open queue is breadth-checked for tickets that agree — then the claim goes to a pass whose job is to refute it.
No human nudge. Nobody was awake to give one.
Twenty-four minutes after the first customer did.
Ordinary tickets look ordinary. Klaxo breadth-checks the open queue for the ones that agree, and only a claim that survives a deliberately skeptical second pass is allowed to raise anything. An unconfirmed one quietly becomes a single engineering ticket instead of waking someone.
checkout-api 5xx 7.4% vs 0.31% baseline
A week of the entity's own behaviour. Not a number somebody guessed a year ago.
Nothing to write. Nothing to tune. The metric is compared against itself.
Four times its own baseline, and still climbing.
A noisier lane needs a stricter gate — so a fire is checked against the queue and challenged by a second pass before it can raise anything.
Coupon service returning 500s on ~8% of carts since 02:58, right after a config push to pricing-svc. Real, customer-facing, not self-healing.
Metric, log and cluster-state rules on a loop, deduped on the firing fingerprint so a restart never doubles an incident.
Describe it in plain language, then see whether it would have caught the outages you remember — before it can page anyone.
No authored number at all. Machine-seeded sweeps land as proposals for a human to arm.
Every ticket resolves to exactly one of four outcomes: reply to the reporter, file an engineering ticket, raise an incident, or name the manual step a human has to take. Drafted, evidenced, waiting on one click.
pricing-svc returns ₹0 GST on 3 pincodes
held
Hi Ritu — your render for LSP-71204 was stuck behind a render-worker
backlog, not a problem with your design. The queue is cleared and yours is second in line.
Expect it in your project by 04:10. I’ll update this ticket the moment it finishes — nothing needed from you, and your SLA credit is already applied.
checkout-api 5xx 7.4% vs 0.31% baseline
raised
A tool-using agent queries logs and metrics, walks live cluster state, reads the recent commits and the runbooks, and returns a verdict — cause, blast radius, whether it is customer-facing. Nobody writes a prompt.
checkout-api 5xx at 7.4% against a 0.31% last-week baseline — 24x, three consecutive 30s windows.
asyncpg.PoolTimeoutError, prod
pricing-svc@8f2c1d config, last 6h
The 03:24 push took pricing-svc from 200 pooled connections down to 40, so the pool saturates at 40/40 and checkout-api blocks waiting for a quote. Not traffic: RPS flat at 1.02x. Not a checkout rollout: zero restarts, no image change in 6h.
#checkout-oncall
SEV-2
revert held
8f2c1d drafted, awaiting approvalAuthored rules on one side, a baseline that needs no threshold on the other. Plus the support queue, which no monitoring tool reads.
Corroborated across the queue, then handed to a pass whose job is to refute it. Only survivors get raised.
Logs, metrics, live cluster state, the recent commits, the runbooks. Nobody writes a prompt.
Raise, reply, file, review, propose a fix. How far it goes alone is a dial you set.
Every outcome becomes a fact the next investigation already knows. That is why the second one is faster.
Four settings. Start at zero, watch what it would have done, move the dial when it has earned it. One deliberate exception: a verified outage is always raised, even when everything else waits for a click.
Similar incidents collapse into one cause, so four tickets read as one problem — and the gaps get attributed. Your retro brief is written before the meeting starts.
18 of 47 are one cause: catalog 5xx spike, every one between 02:00 and 04:30,
every one downstream of the pricing-svc price sync. RCA graded strong; the fix is
still unshipped. Four signatures account for 40 of the 47; the other 7 are one-offs.
No migration. No re-platforming.
A promised turnaround is a commitment somebody has to keep. So it comes out of a table, checked against 704 real replies. When the request doesn't match cleanly, the reply carries no date at all — silence beats promising two days on a ten-day job.
And when they do, the investigation should already be finished.