Skip to content

Finding an incident’s root cause, with evidence you can check

How to use Botify for incident root cause analysis: gather evidence, align one time window, link entities across tools and rank causes you can trace and verify.

Updated · 6 min read

What can AI do in root cause analysis, and what can it not?

Most of the time in a root cause analysis goes into collecting and lining up evidence: opening the monitoring problem, finding the deployments, searching logs, checking the change calendar and working out which host or service each tool is talking about. Botify with API access to those tools can do that collection in parallel and in minutes, and it does not forget to check the second system.

What AI cannot do is prove causation by itself. The Google SRE book describes troubleshooting as the hypothetico-deductive method: form hypotheses from observations, then test them until the cause is found. Botify is very good at the first half and useful for the second, but the final call, and any test that changes production, belongs to an engineer.

So the goal is not an AI that announces the root cause. It is an AI that produces a ranked, sourced list of candidate causes that an engineer can confirm or reject quickly.

Step 1: Gather problems, deployments, logs and changes

Start from the symptom and collect four kinds of evidence for the affected service. Each answers a different question, and skipping one is the most common reason an investigation lands on the wrong cause.

  • Problems and alerts: what broke, when it started, which entities are affected. In Dynatrace this is the problem with its root-cause evidence; in Splunk it may be fired alerts or ITSI notable events.
  • Deployments: what software changed. Deployment events in the observability platform are the fastest source, if your pipeline records them.
  • Logs, traces and metrics: what the system actually did. Look for new error types, latency by operation, and saturation on dependencies.
  • Change records: what people changed that is not a deployment, such as configuration, firewall, database or certificate changes approved in ServiceNow.

Step 2: Align everything on one time window

Before comparing anything, fix one investigation window and apply it to every query: from a lookback before the first symptom (often one to two hours for a sudden failure, longer for slow degradation) to now or to recovery. Every tool call uses the same start and end, in UTC, so results can be compared directly.

Different systems also disagree about time. Systems run on different clocks, some tools stamp events when they are received rather than when they happened, and ITSM records often carry a planned window rather than the real one. Compare events with a small tolerance instead of exact timestamps, and prefer the observed time over the planned time when both exist.

Step 3: Link the same entity across tools

The same service rarely has the same name everywhere. Dynatrace identifies it by an entity ID, Splunk by a host, source or index field, and ServiceNow by a configuration item with its own sys_id. Until those are linked, Botify cannot know that a change on one CI and an error spike on one host are about the same thing.

Treat the mapping as data, not guesswork. Look up the entity by name in each tool, confirm matches using properties such as host names, tags or relationships, and save confirmed links so the next investigation starts with them. A wrong link silently corrupts every later correlation, so a person should confirm new ones.

Step 4: Rank likely causes instead of naming one

With aligned, linked evidence, Botify can build a short list of candidate causes and rank them. Good ranking criteria are the ones an experienced engineer already uses. The table shows the result for an illustrative checkout incident.

  • Time order: the candidate change happened before the first symptom, not after it.
  • Topology: the change touched the failing service or something it depends on.
  • Fit: the kind of change matches the kind of failure, for example a connection-pool setting and a spike in connection timeouts.
  • Blast radius: the affected entities match what the change could reach.
  • History: a similar change caused a similar incident before.
Candidate causeSupporting evidenceEvidence againstNext check
Release 4.12 of the checkout serviceDeployed 6 minutes before the error spike; new timeout errors in logsNone found yetCompare error rate on hosts still on 4.11
Database parameter changeApproved change in the same window on the database CIDatabase latency stayed flatCheck connection-pool metrics
Upstream payment providerSome errors mention the providerErrors started before provider calls roseLow priority unless the first two are ruled out

Step 5: Keep every finding traceable to evidence

An RCA written by AI is only useful if every sentence can be checked. Each finding should point to the tool call that produced it: which query, against which system, for which window, and what it returned. Anything the model inferred without a direct source should be labelled as a hypothesis, not stated as fact.

This also protects long investigations. When Botify summarises large results to fit its context, the full result must still be retrievable, so the final report can quote the original record rather than a paraphrase of a paraphrase.

Common pitfalls: clock skew, correlation vs causation and more

The failure modes of AI-assisted RCA are mostly the classic failure modes of human RCA, repeated faster. Watch for these.

  • Clock skew: comparing exact timestamps across systems makes a cause look like it came after the effect, or the reverse. Use a tolerance and state it in the report.
  • Correlation taken as causation: the SRE book recounts an investigation where a latency increase correlated with one datastore call, and the theory collapsed when unrelated static-content requests turned out to be slow too. A match in time is a lead to test, not a verdict.
  • Anchoring on the last deployment: recent releases are easy to find, so they get blamed first. Check infrastructure and configuration changes with equal weight.
  • Missing data read as absence: “no errors in the logs” may mean the wrong index or a too-narrow window. Botify should say what it searched.
  • Unbounded queries: a search without a time window or a result cap can time out or cost a lot. Bound every query.

How Botify supports root cause analysis

Botify includes platform correlation tools for exactly these steps. correlation.find_entity finds the same service, host or application across connected systems; correlation.link_entity records that an identifier in one tool refers to that entity, and always waits for human approval; correlation.time_window_overlap checks whether two windows overlap, with a tolerance for clocks that disagree; and correlation.evidence_fetch retrieves the full evidence captured earlier in the run so every finding traces back to its source.

Those tools work alongside read-only connector tools for Dynatrace problems and deployments, Splunk searches and ITSI notable events, and ServiceNow incidents and changes near a given time. The result can be delivered as a report in PDF, XLSX, DOCX or CSV, in English or Arabic.

Frequently asked questions

Can AI find the root cause of an incident automatically?

AI can gather evidence and rank likely causes much faster than a person, but it cannot prove causation on its own. Treat its output as ranked hypotheses with evidence, and let an engineer confirm the cause.

What data does AI need for incident root cause analysis?

At minimum: the problem or alert, deployment events, logs or traces, and change records for the affected service and time window. Without change records, AI tends to blame the most visible deployment.

How do you handle clock skew between monitoring tools?

Use one investigation window in UTC for every query and compare events with a tolerance of a few minutes rather than exact timestamps. Prefer observed event times over planned or received times.

Is AI root cause analysis the same as Dynatrace Davis or AIOps RCA?

No. Built-in RCA in an observability platform analyses the data inside that platform. Botify works across platforms and adds change records, logs from other tools and past incidents, often starting from the platform’s own problem as the first clue.

How do you check that an AI root cause analysis is correct?

Every finding should link to the exact query or record that supports it. Reviewers check the evidence, look for evidence against the top candidate, and run a test that could disprove it before accepting it.

Sources

  1. Google SRE Book, Chapter 12: Effective Troubleshooting

Where this applies in Botify

Related articles

All articles

Start with one team.

Pick one workflow and one or two systems. We connect Botify and measure the difference.