Every incident arrives already investigated.

Klaxo watches your metrics, your logs, your cluster — and your support queue. When something breaks it raises the incident itself, and the answer is already written.

Live docket · 63 services, 41 signals/min 03:58 UTC
Docket · 4 open, 4 closed
INC-3312 checkout-api capture 5xx 03:42
SUP-4488 payment page blank after redirect 03:41
INC-3309 pricing-svc p99 over 2s 03:12
INC-3307 cart-worker queue 18.4k 02:56
SWP-118 nightly sweep, 63 services 03:00
INC-3304 identity token refresh 01:38
INC-3301 search-index replica lag 00:52
INC-3298 catalog CDN purge backlog 23:14
INC-3312 · checkout-api · http_5xx_ratio · 02:30–03:58 UTC
5xx ratio now
7.4%
24× baseline
Failed captures
1,208
since 03:41
Time to verdict
94s
no human involved
8.0% 4.0% 0.0% learned baseline corridor · 0.20–0.50% 03:24 · config push · pricing-svc 8f2c1d 03:41 breach declared 03:42:34 verdict written 02:30 03:58
Verdict

Real outage. Raised as INC-3312, sev-2. checkout-api 5xx on POST /v2/capture went from 0.31% to 7.4% in the 17 minutes after the 03:24 config push to pricing-svc. Confined to the payment-capture path — browse and cart are nominal.

Cause is upstream, not checkout-api itself: push 8f2c1d took the pricing-svc pool from 200 connections down to 40, the pool now saturates at 40/40, and checkout-api fails closed instead of falling back to the cached quote. 11 of 14 checkout-api pods report PoolTimeout; the three still holding a warm connection are clean.

Corroborated by the support queue — 3 tickets in 9 minutes from distinct users, all on the same payment step: SUP-4483 opened it at 03:33, then SUP-4485 at 03:37, then SUP-4488 at 03:41 — the ticket linked to this incident. Breadth confirmed, so this was auto-raised rather than held.

Incident raised 03:42 Revert 8f2c1d drafted Pool 40→200 awaiting approval 3 tickets linked Payments on-call paged
Evidence: metrics on 4 services, 812 log lines across 14 pods, 3 tickets  ·  6 deploys in the window, one correlated  ·  verdict written before any human acknowledged it.

Most outages reach a customer before they reach a dashboard.

03:33

A customer writes in.

One ticket, on a night with twenty other one-offs. Nothing about it looks like an outage yet.

03:37 · 03:41

Then two more that rhyme.

Different words. Same shape.

Verification

Corroborated, then challenged.

The open queue is breadth-checked for tickets that agree — then the claim goes to a pass whose job is to refute it.

03:42

It raises the incident itself.

No human nudge. Nobody was awake to give one.

03:57

Monitoring finally notices.

Twenty-four minutes after the first customer did.

the same outage, two sensorssupport queue · 1 signal
SUPPORT QUEUE METRIC 03:3003:35 03:4003:45 03:5003:55 04:00 03:33 03:37 03:41 CORROBORATED ×3 RAISED 03:42 · KLAXO 9 min KLAXO ACTED ALERT THRESHOLD ALERT FIRES 03:57 24 min BEFORE MONITORING KNEW

Three complaints. One cause.

Ordinary tickets look ordinary. Klaxo breadth-checks the open queue for the ones that agree, and only a claim that survives a deliberately skeptical second pass is allowed to raise anything. An unconfirmed one quietly becomes a single engineering ticket instead of waking someone.

support queue — sensor lane03:29–03:44 · 8 new · scanning
Customer tickets, classified on arrival 3 correlated
SUP-4489 Coupon FESTIVE20 not applying pricing-svc 03:44
SUP-4488 Payment page blank after redirect checkout-api 03:41
SUP-4487 Product images never load catalog 03:39
SUP-4486 Password reset email never arrives awaiting user 03:38
SUP-4485 Checkout error, tried 4 times checkout-api 03:37
SUP-4484 Cart empties when I switch tabs auto-replied 03:35
SUP-4483 Card payment stuck, then fails checkout-api 03:33
SUP-4482 Order still says packed order-svc 03:29
correlated SUP-4483 03:33 · SUP-4485 03:37 · SUP-4488 03:41 — 3 tickets in the 9 min to the raise, one payment step
outage raised INC-3312 · sev-2 · 03:42 — checkout-api 5xx at 7.4% vs 0.31% baseline
Baseline

It learns what normal looks like.

A week of the entity's own behaviour. Not a number somebody guessed a year ago.

Armed

The corridor is the threshold.

Nothing to write. Nothing to tune. The metric is compared against itself.

Breakout

Then something leaves it.

Four times its own baseline, and still climbing.

Raised

Verified before anyone is paged.

A noisier lane needs a stricter gate — so a fire is checked against the queue and challenged by a second pass before it can raise anything.

checkout-api · 5xx ratebaseline · learning
BREAKOUT BASELINE CORRIDOR ±2.6σ · TRAILING 7d
Verdict · 4m 27s after the first signal

Coupon service returning 500s on ~8% of carts since 02:58, right after a config push to pricing-svc. Real, customer-facing, not self-healing.

Incident raisedInvestigatingThread opened
Rule lane
Authored thresholds

Metric, log and cluster-state rules on a loop, deduped on the firing fingerprint so a restart never doubles an incident.

Backtest
Replayed over history

Describe it in plain language, then see whether it would have caught the outages you remember — before it can page anyone.

Anomaly lane
Its own last week

No authored number at all. Machine-seeded sweeps land as proposals for a human to arm.

Not a category. The next action.

Every ticket resolves to exactly one of four outcomes: reply to the reporter, file an engineering ticket, raise an incident, or name the manual step a human has to take. Drafted, evidenced, waiting on one click.

Action docket overnight shift · 03:07–03:42 UTC
SUP-4462 Refund not credited 6 days after cancellation 03:07 reply sent
ENG-2149 pricing-svc returns $0 tax on 3 postal codes 03:24 eng ticket held
SUP-4471 Export stuck at Processing for 46 min 03:38 reply held
Drafted reply · SUP-4471
drafted 03:38:12 · job EXP-71204

Hi Dana — your export EXP-71204 was stuck behind an export-worker backlog, not anything wrong on your side. That backlog is draining now and yours is second in the queue.

Expect it by 04:10. I’ll update this ticket the moment it finishes — nothing is needed from you.

queue depth now
24
peak 306 at 03:31
promised eta
04:10
2nd in line
gate score
0.91
above your send threshold
Approve & send Edit draft held 4 min
SUP-4474 Duplicate line items on quote Q-51188 03:41 manual held
INC-3312 checkout-api 5xx 7.4% vs 0.31% baseline 03:42 incident raised
Five tickets, five decisions, none of them a human’s. Two drafted actions wait on approval, one is flagged for manual handling · median time to verdict 41s

It reads production, then commits to an answer.

A tool-using agent queries logs and metrics, walks live cluster state, reads the recent commits and the runbooks, and returns a verdict — cause, blast radius, whether it is customer-facing. Nobody writes a prompt.

Investigation trace · checkout-api 6 tool calls · 03:41:00 → 03:42:34 UTC
Trigger

Rule R-014: checkout-api 5xx above 5% for three consecutive 30s windows.

01 · metrics.query

5xx ratio 0.31% → 7.4% across 12 consecutive windows — 24× the same hour last week.

02 · logs.search

asyncpg.PoolTimeoutError 2,847 times in 9 minutes, on 11 of 14 pods.

03 · cluster.walk

deploy/checkout-api unchanged — 14/14 pods ready, zero restarts in the window.

04 · tickets.scan

3 tickets on the same payment step, 03:33 → 03:41: SUP-4483, SUP-4485, SUP-4488.

05 · commits.diff

pricing-svc pool cap 200 → 40 in 8f2c1d, pushed 03:24 — the one change correlated with the window.

06 · memory.recall

Same signature as INC-2980 on 19 Jun, which was fixed by reverting the push.

Verdict

The 03:24 push cut pricing-svc’s pool cap from 200 to 40. It saturates at 40/40, so checkout-api blocks waiting on a quote and returns 5xx. Not traffic: RPS flat at 1.02x. Not a checkout rollout: zero restarts, no image change in 6h.

Time to verdict
94s
Gate score
0.91
INC-3312 Real outage: pool starvation · sev-2 · paged #checkout-oncall RAISED 03:42 revert held
Nothing above waited on a human. The drafted revert of 8f2c1d is held for approval — config is never pushed autonomously.
01 · DETECT

Two lanes. One sink.

Authored rules on one side, a baseline that needs no threshold on the other. Plus the support queue, which no monitoring tool reads.

02 · TRIAGE

Real, or not.

Corroborated across the queue, then handed to a pass whose job is to refute it. Only survivors get raised.

03 · INVESTIGATE

It reads production, then commits.

Logs, metrics, live cluster state, the recent commits, the runbooks. Nobody writes a prompt.

04 · ACT

Drafted first. Executed second.

Raise, reply, file, review, propose a fix. How far it goes alone is a dial you set.

05 · LEARN

So it happens once.

Every outcome becomes a fact the next investigation already knows. That is why the second one is faster.

01 Detect metrics, logs, tickets 02 Triage real, or noise 03 Investigate writes the verdict 04 Act raise, reply, propose 05 Learn baselines, playbooks VERDICTS RETUNE DETECTION

You decide how much it does alone.

Four settings. Start at zero, watch what it would have done, move the dial when it has earned it. One deliberate exception: a verified outage is always raised, even when everything else waits for a click.

Autonomy what Klaxo may do without a human current · REVIEW
Reply to the customerinternal note + reporter draft
waits for a click
File an engineering ticketone record, not an outage
waits for a click
Step outside the systemrefund, access, account change
waits for a click
Verified production outagecorroborated, then challenged
RAISED AUTOMATICALLY

What keeps breaking, and which RCAs never got written.

Similar incidents collapse into one cause, so four tickets read as one problem — and the gaps get attributed. Your retro brief is written before the meeting starts.

Incident review 30d · 47 incidents · prod-east-1
Clusters by hour and time-to-recover
60 45 30 15 0 00:00 06:00 12:00 18:00 24:00 mttr target 20m catalog 5xx spike ×18 nodepool disk pressure ×6 cart-worker OOMKilled ×11 pricing-svc pool starve ×5
x hour of day · y mttr min 7 one-offs densest ringed
Recurring causes and RCA grade
catalog 5xx spike · rca strong 18
cart-worker OOMKilled · rca weak 11
nodepool disk pressure · rca adequate 6
pricing-svc pool starve · rca strong 5

18 of the 47 are one signature — catalog 5xx spike, every one between 02:00 and 04:30, every one downstream of the pricing-svc price sync. RCA graded strong; the fix is still unshipped. Four signatures cover 40 of the 47; the other 7 never recurred.

These 47 are real — one production instance, 30 days, every incident Klaxo raised or was handed. The outage walked through earlier on this page is a composite drawn from them.

repeat rate
85%
median mttr · all 47
24m
rca strong
23/47
densest signature SIG-0114 catalog 5xx spike — 18 occurrences · median mttr 26m · last fired 6h ago

It plugs into the stack you already run.

No migration. No re-platforming.

Connected sources prod-east-1 polled every 15s · last sync 12s ago
METRICS victoria-metrics LOGS loki · 62k lines/min CLUSTER gke · 214 pods · 9 nodes WAREHOUSE bigquery · lag 41m TICKETS zendesk · 37 open CHAT slack · #inc-war-room GIT github · main a7c1283 IDENTITY okta scim · 412 users KLAXO 8 sources · 3.1k ev/s verdict p50 41s
2 live streams 7 nominal 1 degraded · bigquery lag 41m
Every source is read over its own vendor API with read-only credentials — no agent, sidecar or DaemonSet on your services.
Observability
  • Prometheus
  • Grafana
  • Loki
  • VictoriaMetrics
  • OpenTelemetry
  • Datadog
  • New Relic
Runtime & cloud
  • Kubernetes
  • Google GKE
  • Amazon EKS
  • Azure AKS
  • Terraform
  • Helm
  • Argo CD
Tickets & ITSM
  • Jira
  • Jira Service Management
  • Linear
  • ServiceNow
  • Zendesk
  • Freshdesk
  • Intercom
Chat & on-call
  • Slack
  • Microsoft Teams
  • Google Chat
  • PagerDuty
  • Opsgenie
  • Webhooks
  • Email
Source & CI
  • GitHub
  • GitLab
  • Bitbucket
  • GitHub Actions
  • Jenkins
  • CircleCI
  • Buildkite
Identity
  • Okta
  • Microsoft Entra ID
  • Google Workspace
  • Keycloak
  • OIDC
  • SAML 2.0
  • SCIM
Data
  • PostgreSQL
  • pgvector
  • BigQuery
  • Snowflake
  • ClickHouse
  • S3 / GCS
  • Kafka
Knowledge
  • Outline
  • Confluence
  • Notion
  • Google Drive
  • Markdown runbooks
  • Backstage
  • OpenAPI specs
Connected today. Everything else speaks a protocol Klaxo already reads — a new source is a config entry, not a release.
PromQL · LogQL · OTLP · Kubernetes API · Git · REST · GraphQL · webhooks · SQL · MCP

The date on a customer reply is looked up. Never generated.

A promised turnaround is a commitment somebody has to keep. So it comes out of a table, checked against the replies your team has already sent. When the request doesn't match cleanly, the reply carries no date at all — silence beats promising two days on something that takes ten.

your turnaround tabledeterministic
bulk record update10 business days
new entity onboarding7 business days
contact or email change5 business days
everything else, matched2 business days
no confident matchno date sent
Your rows, your commitments. Most specific first; first match wins — a long-turnaround request can never be swallowed by a shorter default.

In your cluster. Your data stays there.

Deployment
Your infrastructureKlaxo’s control plane in your VPC — nothing added to your services
Data
Never leavesembeddings and memory on your own database
Access
Read-only by defaultwrite scopes granted one capability at a time
Models
Claude on Vertex AIyour own Google Cloud project — no vendor sees production data

Nobody should hear about an outage from a customer.

And when they do, the investigation should already be finished.