Skip to content

Answers from your sources vs a general model

Why a general model invents facts about your company, how grounding answers in your documents and systems, its limits, and how to test for it.

Updated · 5 min read

What does “grounded” mean for an assistant?

Grounding means tying every answer to evidence the system retrieved for this specific question, rather than to whatever the model absorbed during training. The evidence can be a passage from a policy document, a record from a ticketing system, or a metric from a monitoring tool.

A grounded agent has three habits. It looks things up before answering, it answers from what it found and can show where that came from, and it abstains when the sources do not cover the question. The third habit is the one most demos skip and most production incidents trace back to.

Why does a general model make things up about your company?

A general model has never seen your refund policy, your price list or last night’s incident. Asked about them, it does what it was trained to do: produce the most plausible continuation. Plausible and correct are different things, and the model has no internal signal that tells it which one it is producing.

The consequences are real. In February 2024, a Canadian tribunal held Air Canada responsible after its website chatbot told a customer that a bereavement fare could be claimed retroactively, which contradicted the airline’s actual policy. The tribunal rejected the argument that the chatbot was responsible for its own actions, noting that the airline is responsible for all the information on its website, whether it comes from a static page or a chatbot.

The lesson is not that chatbots are dangerous. It is that anything speaking for your company must speak from your company’s sources, and that a wrong answer delivered politely is still your answer.

Grounded assistant vs general model: side by side

General models are excellent tools; they are just the wrong tool for stating facts about your organisation. The table shows where each approach wins.

Note that a grounded agent still uses a general model inside it. The model does the language work, understanding the question, reading the evidence, summarising and translating it, while the sources decide what is true. Grounding changes where facts come from, not which model you use.

CriterionGeneral LLMGrounded assistant
Source of the answerPatterns in training dataYour documents and systems, retrieved at question time
Facts specific to your companyGuessed, often plausiblyRetrieved, or declined if absent
FreshnessFixed at the training cutoffAs current as the source it reads
When no answer existsProduces one anywaySays it does not know, or hands off to a person
AuditabilityNo traceable sourceEach answer traceable to a document or tool result
Updating a factWait for a new model or fine-tuneEdit the source document or fix the system record
General knowledge and writingBroad and strongDeliberately bounded to your domain
Setup effortNoneSources, retrieval, tool access and testing
Best forDrafting, brainstorming, translation, coding helpCustomer answers, operations, anything with liability

How do you ground an assistant?

There are three main techniques, and serious deployments combine them. The right mix depends on where your truth lives: in documents, in operational systems, or in structured databases.

  • Retrieval over documents (RAG): index policies, manuals and runbooks, retrieve the passages relevant to the question and give them to the model as context.
  • Tool calls to live systems: let the assistant query the ticketing, monitoring or CRM system directly, so answers reflect the current state rather than a copy.
  • Structured queries: for numbers such as order status or stock levels, fetch the exact field instead of asking a model to recall or estimate it.
  • Instructions and checks around all three: answer only from provided evidence, cite it, and abstain or escalate when evidence is missing or conflicting.

Is grounding a guarantee against hallucinations?

No. Grounding moves the problem from “the model does not know” to “did the system find and use the right evidence?”, which is far easier to observe and fix, but it can still fail.

  • Retrieval misses: the right passage exists but is not retrieved, so the model answers from a weaker one.
  • Stale or conflicting sources: two versions of a policy disagree and the assistant picks the old one.
  • Ignoring the context: the model blends retrieved facts with its own assumptions.
  • Poisoned content: a document or email contains instructions aimed at the model; treat retrieved text as data, never as commands.
  • Over-reach: the assistant answers a question adjacent to its sources instead of declining.

How to test whether an assistant is really grounded

Build a small test set before launch and rerun it after every change to prompts, sources or models. Twenty or thirty well-chosen questions catch most regressions.

  • Questions with known answers in your sources: the assistant should answer correctly and point to the source.
  • Questions with no answer in your sources: the assistant should say so, not improvise.
  • Questions where two sources conflict: the assistant should flag the conflict or prefer the source you designated as authoritative.
  • A fact you just changed: the assistant should give the new answer immediately.
  • Arabic and English phrasings of the same question: the answers should agree.

Grounding answers in live system data

For operations work, the most reliable ground truth is the system of record itself. Botify grounds answers by calling read-only tools such as servicenow.search_knowledge, dynatrace.get_problem or splunk.run_saved_search, and records every tool call and result in an audit trail, so anyone reviewing an answer can see exactly what it was built on.

Whatever platform you use, apply the same principle: the assistant should be able to show its evidence, and a reviewer should be able to check it without asking again.

Frequently asked questions

Can a grounded agent still answer general questions?

You can allow it, but it is usually better to redirect politely to the agent’s domain. If general answers are allowed, label them clearly so users can tell a company fact from general knowledge.

Is grounding the same as RAG?

RAG, retrieval-augmented generation, is one way to ground a model, using retrieved document passages. Grounding is the broader goal and also covers live tool calls, structured database queries and the rules that force the assistant to cite or abstain.

Does fine-tuning make a model grounded?

Not by itself. Fine-tuning changes style and behaviour and can teach domain vocabulary, but facts learned that way are still recalled from memory, cannot be traced to a source and go stale. Use retrieval or tools for facts, and fine-tuning, if at all, for form.

Does grounding eliminate hallucinations?

It greatly reduces them and makes the remaining ones visible, because each answer can be checked against its evidence. It does not eliminate them: retrieval can miss, sources can be wrong, and models can still blend in assumptions, so testing and review remain necessary.

How do I update what a grounded agent knows?

Update the source: edit the document, fix the system record or add the missing policy. The next question retrieves the new version, with no retraining. This is also why source hygiene, one current version of each policy, matters so much.

Sources

  1. Air Canada found liable for chatbot’s bad advice on plane tickets (CBC News, February 2024)

Where this applies in Botify

Related articles

All articles

Start with one team.

Pick one workflow and one or two systems. We connect Botify and measure the difference.