Change Intelligence Platform

Know why production changed in minutes, not hours

Your infrastructure constantly changes. Deployments, Kubernetes rollouts, host configuration, metric baselines. We collect every change across your stack — and when an incident hits, show the likely causes ranked, with the evidence behind each one.

Self-hosted — data stays in your infra
No eBPF — read-only sensors, any VM
5 min to first events via webhook
whatchanged.io/dashboard

Last 24h

Change summary

43 changes
2 high risk
Metric drift

error_rate p950.4% → 2.1%

Above baseline
DeploymentRollout

billing-apiMemory +28%

After rollout
Host changeChange

api-17vm.swappiness 60 → 10

sysctl diff

Every incident starts with the same question

Answering "what changed?" by hand means digging through Grafana, GitHub, kubectl and Slack. WhatChanged has the timeline and ranked causes ready the moment the alert fires.

BeforeManual investigation
13:00
Alert fires
13:05
Open Grafana
13:20
Check Kubernetes
13:40
Search GitHub commits
13:55
Review Terraform changes
14:10
Found root cause
Total time
70 min
AfterWith WhatChanged
13:00
Alert fires — incident opens
13:00
Timeline + probable causes ready
Probable causes, ranked
Deployment billing-api v2.3.10.87
sysctl changed on api-170.54
ConfigMap billing-config changed0.32

4m before the alert • metrics shifted after it: memory +28%

Total time
1 min
Findings, not noise

Catch drift before it pages you

Every host and metric gets a rolling baseline. When a value drifts 2σ, 3σ or 4σ away, you get a finding — and it resolves itself when things return to normal.

Finding: metric drift

error_rate p95 · billing-api

3σ deviation

Custom Prometheus metric drifted from its rolling baseline — no alert rule needed.

Baseline (12 windows)
0.4%
Current window
2.1%
Per-entity baselineAny PromQL metricAuto-resolves

Code & Delivery

GitHub webhooks turn merges, pushes and pipelines into change events, mapped to services

PR #142 merged
Pipeline failed
Force push to main

Kubernetes

Read-only watchers catch rollouts, config changes, OOM kills, crash loops and node events

Image bump billing-api
ConfigMap changed
OOMKilled ×3

Host Internals

A 6MB agent diffs sysctl, kernel, packages, services and network config — no eBPF

vm.swappiness 60 → 10
Kernel upgraded
openssl upgraded

Metric Drift

Your own Prometheus metrics — error rate, RPS, latency percentiles — against rolling baselines

error_rate p95 +3σ
RPS baseline shift
Any PromQL expression

Alerts → Incidents

An Alertmanager webhook opens an incident with a two-hour change timeline around it

alert.fired opens incident
Timeline T-2h
Auto-resolve on recovery

Ranked Causes

Every candidate scored by timing, scope, metric shifts and change-type risk — fully explainable

Deploy ranked #1 (0.87)
Metrics shifted after it
Same host as incident

Works with your stack, not instead of it

Keep Grafana, Prometheus and your alerting. WhatChanged adds the layer they are missing: what changed, and which change is the likely cause.

Every change class
From a merged PR to a sysctl flip on a host: git, CI, Kubernetes config, packages, kernels, custom metrics — one feed.
No eBPF, no sidecars
Read-only sensors: a single static Go binary per host, watch-only Kubernetes RBAC, PromQL queries. Runs on any VM.
Self-hosted
Your telemetry never leaves your infrastructure. Config values stay in-cluster — sensors ship names and hashes, not secrets.
Explainable, not magic
Every ranked cause shows its score breakdown: how recent, how related, what metrics shifted. No black-box verdicts.

Stop asking "What changed?"

Connect a GitHub webhook and see your first change events in minutes — on your own infrastructure.

Early access • Free while in beta • Self-hosted