IT operations
IT operations through conversation: triage, investigate, report
What Botify does in IT operations: triage alerts, investigate across observability and ITSM tools, write reports, and act only with approval.
Updated · 6 min read
What does Botify do in IT operations?
In IT operations, Botify is an assistant that can call tools: it reads a question or an alert, decides which system to query, calls that system’s API, reads the result and decides what to ask next. The tools are the ones your team already uses, such as an observability platform (Dynatrace, Splunk), an ITSM platform (ServiceNow) and email.
The difference from a chatbot is that Botify does not answer from memory. Every statement it makes about your environment should come from a tool call it made during that run, so the answer can be checked. A chatbot that was trained on public documentation can explain what a Dynatrace problem is; Botify can tell you which problems are open on your checkout service right now and what changed before they started.
In practice Botify works like a junior on-call engineer who never gets tired of copying IDs between tabs: it gathers the evidence, lines it up, and hands a senior engineer a short, sourced summary to decide on.
How is Botify different from classic AIOps?
Classic AIOps tools were built to reduce alert noise. They ingest large volumes of events and metrics, learn baselines, detect anomalies, group related alerts into one incident and, in the better products, suggest a probable root cause from topology. That job is still necessary and those tools do it well.
Botify sits one layer above. It does not replace anomaly detection; it consumes its output. A Dynatrace problem or a Splunk ITSI notable event is a starting point for Botify, which then pulls deployments, logs, change records and past incidents from other tools that the AIOps engine never sees.
| Classic AIOps | Botify | |
|---|---|---|
| Main job | Detect anomalies, reduce noise, group events | Answer questions and run investigations across tools |
| Input | Streams of metrics, events and logs in one platform | A question, an alert or an incident, plus API access to many tools |
| Output | Alerts, problems, episodes, anomaly scores | A written answer or report that cites the evidence it gathered |
| New kinds of question | Needs a new rule, model or dashboard | Handled if the right tools are connected |
| Taking action | Usually runbook automation configured in advance | Proposes an action; a person approves before it runs |
| Typical failure | Noisy or missed alerts when baselines drift | Confident but unsupported conclusions if evidence is not enforced |
What does Botify actually do in IT operations?
Strip away the marketing and there are four jobs Botify can do reliably today. Each one is mostly reading and summarising, which is why they are safe to start with.
- Triage: read a new alert or problem, check whether a related incident already exists in ServiceNow, look at the affected service’s recent changes and say how serious it looks and who probably owns it.
- Investigation: collect problems, deployments, logs, traces and change records for the same service and time window, then rank the likely causes with the evidence for each.
- Reporting: turn an investigation into a shift handover, an incident summary or a weekly problem review, in the language and format the reader needs.
- Safe actions: draft the incident update, the stakeholder email or the deployment marker, then pause until an authorised person approves it.
What does an AI-led investigation look like end to end?
Here is a realistic flow for the question “Why is checkout slow since 10:00?”, using tools that exist in production connectors today. The order is not scripted; the model chooses the next call from what the previous one returned.
- List open problems in Dynatrace for the last two hours and fetch the one that affects the checkout service, including its root-cause evidence and affected entities.
- List deployment and change events in Dynatrace for the same window, and change requests near that time in ServiceNow.
- Run a bounded Splunk search for errors on the checkout hosts in the same window, and check whether any ITSI notable events fired.
- Resolve the service name across tools, because Dynatrace, Splunk and ServiceNow rarely call it the same thing.
- Check which changes overlap the start of the problem in time, allowing for clocks that disagree by a few minutes.
- Write a short answer: the most likely cause, the evidence behind it, what would disprove it, and a proposed next step that waits for approval.
Which actions should Botify be allowed to take?
The useful rule is to classify every tool by what it can change, not by how clever the model is. Reading a metric is harmless; posting a deployment event, updating an incident or emailing a customer is not. Most teams start with Botify fully read-only and add write tools one at a time, each behind an approval.
Botify labels every tool with a risk level (read, write, destructive or external send) and an approval mode. Its ServiceNow tools for creating an incident or changing its state, and its Dynatrace tool for recording a deployment event, always wait for a person to approve, and every request, tool call and approval is kept in an audit trail.
| Action class | Example | Sensible default |
|---|---|---|
| Read | Query metrics, search logs, list problems or change records | Runs directly |
| Write | Update an incident state, record a deployment event | Human approval |
| Destructive | Delete, restart, roll back | Not connected at first; approval if ever enabled |
| External send | Email or call someone outside the team | Human approval |
Where do Botify for IT operations fail?
Assistants fail in predictable ways, and each failure has a design answer. None of them is solved by a better prompt alone.
- Invented causes: the model fills a gap with a plausible story. Require every finding to cite a tool result, and flag anything uncited as a hypothesis.
- Runaway queries: an unbounded log search can be slow and expensive. Enforce a mandatory time window and a result cap on every search tool.
- Name mismatch: “checkout”, “checkout-svc-prod” and a ServiceNow configuration item may be the same thing. Keep a mapping of entity identifiers across tools, and let a person confirm new links.
- Correlation taken as causation: a deployment that happened near the incident is a lead, not a verdict. The Google SRE book describes troubleshooting as forming hypotheses and testing them, and that discipline applies to assistants too.
- Too much access: Botify inherits whatever the connected token can do. Grant read scopes by default and add write scopes deliberately.
How to start with Botify in IT operations
Start narrow. Pick one service with good telemetry and an on-call team willing to compare Botify’s answers with their own, connect the observability and ITSM tools read-only, and list the ten questions your engineers ask most during incidents.
Judge Botify on whether its answers are correct and whether every claim can be traced to its source, not on how fluent it sounds. When the read-only answers are trusted, add one write action, such as drafting the incident update, behind an approval, and review the audit trail after each incident.
Frequently asked questions
Does Botify replace AIOps platforms?
No. AIOps platforms detect anomalies and group events; Botify uses those detections as starting points and investigates across other tools. Most teams keep their AIOps platform and add Botify on top.
Can Botify fix incidents automatically?
It can propose a fix and prepare the action, but in a well-governed setup any change waits for a person to approve it. Fully automatic remediation is better left to deterministic runbooks with tested rollbacks.
What data does Botify for IT operations need?
API access to the tools engineers already use during incidents: observability (problems, metrics, logs, traces, deployments), ITSM (incidents, changes, configuration items) and communication channels. It does not need its own copy of your telemetry.
How do you stop Botify from hallucinating a root cause?
Make evidence mandatory: every finding must point to a specific tool result, and anything without one is labelled as a hypothesis. Engineers should be able to open the underlying query or record from the report.
Is Botify safe to connect to production monitoring?
Read-only access to monitoring is low risk if Botify uses scoped credentials and bounded queries. The risk comes from write tools, which should be off by default and approval-gated when enabled.