Scenario
You are operating a production observability platform. The following symptoms appear:
- A new metric was added
- Prometheus memory has spiked
Available evidence:
- prometheus_tsdb_head_series climbs
- prometheus memory alerts
Your task
Determine the cause, recover, document, and validate.
Investigation
The investigation follows the discipline taught in Part XCVIII:
- Form hypothesis, find evidence, test, validate.
- Use the available evidence above to bound the search.
- Reach one of the likely root causes.
Recovery procedure
(Do not reveal until you have reasoned through the problem.)
- Identify the failing component.
- Apply the remediation pathway.
- Validate with the verification step.
- Document the incident.
Remediation
- Identify the offending metric (tsdb). 2. Drop the label via relabel. 3. Restart Prometheus. 4. Validate.
Verification
Memory returns to baseline; offending metric is filtered.
Rollback
Revert relabel change
Prevention
CI lint rules for high-cardinality labels; quarterly review.