Every incident arrives already investigated.

Klaxo watches your metrics, your logs, your cluster — and your support queue. When something breaks it raises the incident itself, and the answer is already written.

Klaxo — Incident Intelligence 41 signals/min · watching 63 services · 03:58 IST
Open docket
INC-3312 checkout-api 5xx at capture 03:41
INC-3309 pricing-svc p99 over 2s 03:12
INC-3307 cart-worker queue 18.4k 02:56
LSD-4488 payment page blank on UPI 03:41
INC-3304 identity token refresh 01:38
INC-3301 search-index replica lag 00:52
INC-3298 catalog CDN purge backlog 23:14
SWP-118 nightly sweep, 63 services 03:00
INC-3312 · checkout-api · http_5xx_ratio, 90 min window
Error ratio now
7.4%
baseline 0.31%
Failed captures
1,208
since 03:41
Time to verdict
94s
no human involved
8.0% 6.0% 4.0% 2.0% 0.0% learned baseline corridor · 0.18–0.52% 03:24 config push pricing-svc 8f2c1d 03:41 breach · 24x baseline verdict written 03:42:34 02:30 02:52 03:14 03:36 03:58
Verdict

Real outage. Raised as INC-3312, sev-2. checkout-api 5xx on POST /v2/capture went from 0.31% to 7.4% in the 17 minutes after the 03:24 config push to pricing-svc. Confined to the payment-capture path — browse and cart are nominal.

Cause is upstream, not checkout-api itself: push 8f2c1d took the pricing-svc pool from 200 connections down to 40, the pool now saturates at 40/40, and checkout-api fails closed instead of falling back to the cached quote. 11 of 14 checkout-api pods report PoolTimeout; the three still holding a warm connection are clean.

Corroborated by the support queue — 3 tickets in 9 minutes from distinct users, all on the same payment step (LSD-4488 leading). Breadth confirmed, so this was auto-raised rather than held.

Incident raised 03:42 Revert 8f2c1d drafted Pool 40→200 awaiting approval 3 tickets linked Payments on-call paged
INC-3312 · evidence: 3 dashboards, 812 log lines, 14 pods, 3 tickets  ·  correlated config push 8f2c1d  ·  verdict written 03:42:34, before first human ack  ·  last sweep 03:00 IST
The gap

Most outages reach a customer before they reach a dashboard.

03:14:02

A customer writes in.

One ticket, on a night with twenty other one-offs. Nothing about it looks like an outage yet.

03:21 · 03:23

Then two more that rhyme.

Different words. Same shape.

Verification

Corroborated, then challenged.

The open queue is breadth-checked for tickets that agree — then the claim goes to a pass whose job is to refute it.

03:23:09

It raises the incident itself.

No human nudge. Nobody was awake to give one.

03:38

Monitoring finally notices.

Twenty-four minutes after the first customer did.

the same outage, two sensorssupport queue · 1 signal
03:10 03:15 03:20 03:25 03:32 03:40 SUPPORT QUEUE METRIC CORROBORATED INCIDENT RAISED ALERT THRESHOLD ALERT FIRES 03:38 24 min UNDETECTED
The queue is a sensor

Three complaints. One cause.

Ordinary tickets look ordinary. Klaxo breadth-checks the open queue for the ones that agree, and only a claim that survives a deliberately skeptical second pass is allowed to raise anything. An unconfirmed one quietly becomes a single engineering ticket instead of waking someone.

support queue — sensor lane03:29–03:44 · 8 new · scanning
Customer tickets, classified on arrival 3 correlated
LSD-4489 Coupon FESTIVE20 not applying pricing-svc 03:44
LSD-4488 Payment page blank after UPI checkout-api 03:41
LSD-4487 Product images never load catalog 03:39
LSD-4486 Password reset email never arrives awaiting user 03:38
LSD-4485 Checkout error, tried 4 times checkout-api 03:37
LSD-4484 Cart empties when I switch tabs auto-replied 03:35
LSD-4483 Card payment stuck, then fails checkout-api 03:32
LSD-4482 Order still says packed cart-worker 03:29
correlated LSD-4488 · LSD-4485 · LSD-4483 — 3 tickets, 9 min, one payment step
outage raised INC-3312 sev-2 03:42checkout-api 5xx 7.4% vs 0.31% baseline
Baseline

It learns what normal looks like.

A week of the entity's own behaviour. Not a number somebody guessed a year ago.

Armed

The corridor is the threshold.

Nothing to write. Nothing to tune. The metric is compared against itself.

Breakout

Then something leaves it.

Four times its own baseline, and still climbing.

Raised

Verified before anyone is paged.

A noisier lane needs a stricter gate — so a fire is checked against the queue and challenged by a second pass before it can raise anything.

checkout-api · 5xx ratebaseline · learning
BREAKOUT BASELINE ±2.6σ · TRAILING 7d
Verdict · 4m 27s after the first signal

Coupon service returning 500s on ~8% of carts since 02:58, right after a config push to pricing-svc. Real, customer-facing, not self-healing.

Incident raisedInvestigatingThread opened
Rule lane
Authored thresholds

Metric, log and cluster-state rules on a loop, deduped on the firing fingerprint so a restart never doubles an incident.

Backtest
Replayed over history

Describe it in plain language, then see whether it would have caught the outages you remember — before it can page anyone.

Anomaly lane
Its own last week

No authored number at all. Machine-seeded sweeps land as proposals for a human to arm.

Triage

Not a category. The next action.

Every ticket resolves to exactly one of four outcomes: reply to the reporter, file an engineering ticket, raise an incident, or name the manual step a human has to take. Drafted, evidenced, waiting on one click.

Action docket 5 actions · every one already decided · 3 held for approval
Decided actions 03:07 – 03:42 IST
LSD-4462 reply Refund not credited 6 days after cancellation 03:07 sent
ENG-2149 eng ticket pricing-svc returns ₹0 GST on 3 pincodes 03:24 held
LSD-4471 reply 3D render stuck at Processing for 46 min 03:38 held
Drafted reply · LSD-4471 · project LSP-71204
drafted 03:38:12 · evidence ENG-2153 · render-worker queue 306 → 24

Hi Ritu — your render for LSP-71204 was stuck behind a render-worker backlog, not a problem with your design. The queue is cleared and yours is second in line.

Expect it in your project by 04:10. I’ll update this ticket the moment it finishes — nothing needed from you, and your SLA credit is already applied.

queue depth now
24
peak 306 at 03:31
render eta
04:10
2nd in line
confidence
0.91
evidence ENG-2153
Approve & send Edit draft held 4 min
LSD-4474 manual Duplicate line items on quote Q-51188 03:41 held
INC-3312 incident checkout-api 5xx 7.4% vs 0.31% baseline 03:42 raised
5 of 5 actions decided without a human · 3 waiting on one click median time to verdict 41s
Investigate

It reads production, then commits to an answer.

A tool-using agent queries logs and metrics, walks live cluster state, reads the recent commits and the runbooks, and returns a verdict — cause, blast radius, whether it is customer-facing. Nobody writes a prompt.

Investigation trace · checkout-api autonomous · 6 tool calls · 94s
Trigger 03:41:00 IST · rule R-014

checkout-api 5xx at 7.4% against a 0.31% last-week baseline — 24x, three consecutive 30s windows.

01 metrics.query checkout-api 5xx · 12×30s windows 24x baseline
02 logs.search asyncpg.PoolTimeoutError, prod 2,847 in 9m
03 cluster.walk deploy/checkout-api · 14/14 ready restarts 0 · HPA idle
04 tickets.scan LSD queue — payment page blank 3 tickets · 03:32–03:41
05 commits.diff pricing-svc@8f2c1d config, last 6h PG_POOL_MAX 200→40 · 03:24
06 memory.recall same signature, prior incident INC-2980 · 19 Jun
Verdict
INC-3312 Real outage — pool starvation RAISED 03:42:34

The 03:24 push took pricing-svc from 200 pooled connections down to 40, so the pool saturates at 40/40 and checkout-api blocks waiting for a quote. Not traffic: RPS flat at 1.02x. Not a checkout rollout: zero restarts, no image change in 6h.

Time to verdict
94s
Confidence
0.94
5xx now
7.4%
DISPATCH owner #checkout-oncall SEV-2 revert held
Verdict written before first human ack · revert of 8f2c1d drafted, awaiting approval
01 · DETECT

Two lanes. One sink.

Authored rules on one side, a baseline that needs no threshold on the other. Plus the support queue, which no monitoring tool reads.

02 · TRIAGE

Real, or not.

Corroborated across the queue, then handed to a pass whose job is to refute it. Only survivors get raised.

03 · INVESTIGATE

It reads production, then commits.

Logs, metrics, live cluster state, the recent commits, the runbooks. Nobody writes a prompt.

04 · ACT

Drafted first. Executed second.

Raise, reply, file, review, propose a fix. How far it goes alone is a dial you set.

05 · LEARN

So it happens once.

Every outcome becomes a fact the next investigation already knows. That is why the second one is faster.

Detect rules · anomalies · tickets Triage real, or not Investigate writes a verdict Act raise · reply · propose Learn remembered FEEDS THE NEXT ONE
Control

You decide how much it does alone.

Four settings. Start at zero, watch what it would have done, move the dial when it has earned it. One deliberate exception: a verified outage is always raised, even when everything else waits for a click.

autonomycurrent · REVIEW
Reply to the customerinternal note + reporter draft
File an engineering ticketone record, not an outage
Step outside the systemrefund, access, account change
Verified production outagecorroborated, then challenged
Learn

What keeps breaking, and which RCAs never got written.

Similar incidents collapse into one cause, so four tickets read as one problem — and the gaps get attributed. Your retro brief is written before the meeting starts.

Incident review 30d · 47 incidents · gke-keystone
Clusters by hour and time-to-recover
60 45 30 15 0 00:00 06:00 12:00 18:00 23:59 mttr target 20m catalog 5xx spike · 18 nodepool disk pressure · 6 cart-worker OOMKilled · 11 pricing-svc pool starve · 5
x hour of day · y mttr min densest ringed
Recurring causes · 30d
catalog 5xx spike 18 strong
cart-worker OOMKilled 11 weak
nodepool disk pressure 6 adequate
pricing-svc pool starve 5 strong

18 of 47 are one cause: catalog 5xx spike, every one between 02:00 and 04:30, every one downstream of the pricing-svc price sync. RCA graded strong; the fix is still unshipped. Four signatures account for 40 of the 47; the other 7 are one-offs.

repeat rate
85%
median mttr
26m
rca strong
23/47
densest signature SIG-0114 catalog 5xx spike 18 occurrences · median mttr 26m · last fired Aug 07 03:12

It plugs into the stack you already run.

No migration. No re-platforming.

Connected sources gke-keystone handshake 03:52:07 · poll 15s
METRICS victoria-metrics LOGS loki · 62k lines/min CLUSTER 214 pods · 9 nodes WAREHOUSE bigquery · lag 41m TICKETS zendesk · 37 open CHAT slack #inc-war-room GIT infra-flux · a7c1283 IDENTITY okta scim · 412 users KLAXO 8 sources · 3.1k ev/s verdict p50 41s
2 live streams 7 nominal 1 degraded warehouse backfill held for approval read-only credentials · nothing installed in your cluster
Observability
  • Prometheus
  • Grafana
  • Loki
  • OpenTelemetry
  • Datadog
  • New Relic
  • Elastic
Runtime & cloud
  • Kubernetes
  • Amazon EKS
  • Google GKE
  • Azure AKS
  • Terraform
  • Helm
  • Argo CD
Tickets & ITSM
  • Jira
  • Jira Service Management
  • Linear
  • ServiceNow
  • Zendesk
  • Freshdesk
  • Intercom
Chat & on-call
  • Slack
  • Microsoft Teams
  • Google Chat
  • PagerDuty
  • Opsgenie
  • Webhooks
  • Email
Source & CI
  • GitHub
  • GitLab
  • Bitbucket
  • GitHub Actions
  • Jenkins
  • CircleCI
  • Buildkite
Identity
  • Okta
  • Microsoft Entra ID
  • Google Workspace
  • Keycloak
  • OIDC
  • SAML 2.0
  • SCIM
Data
  • PostgreSQL
  • pgvector
  • BigQuery
  • Snowflake
  • ClickHouse
  • S3 / GCS
  • Kafka
Knowledge
  • Confluence
  • Notion
  • Outline
  • Google Drive
  • Markdown runbooks
  • Backstage
  • OpenAPI specs
If it speaks an API, Klaxo can speak to it.
PromQL · LogQL · OTLP · Kubernetes API · Git · REST · GraphQL · webhooks · SQL · MCP
Determinism

The date on a customer reply is looked up. Never generated.

A promised turnaround is a commitment somebody has to keep. So it comes out of a table, checked against 704 real replies. When the request doesn't match cleanly, the reply carries no date at all — silence beats promising two days on a ten-day job.

turnaround lookupdeterministic
bulk record update10 business days
new entity onboarding7 business days
contact or email change2 business weeks
everything else, matched2 business days
no confident matchno date sent
Most specific first; first match wins. A long-turnaround request can never be swallowed by the two-day default.

In your cluster. Your data stays there.

Deployment
Your infrastructureyour VPC, your credentials, your egress rules
Data
Never leavesembeddings and memory on your own database
Access
Read-only by defaultwrite scopes granted one capability at a time
Models
Hosted where you areno third party sees your production data
Klaxo — 15 service baselines, 03:12–04:04 IST: checkout-api the only trace off its own corridor

Nobody should hear about an outage from a customer.

And when they do, the investigation should already be finished.