Concepts
Human approval before an assistant acts
How approval gates pause an assistant before risky actions: approval modes, risk levels, what approvers see, timeouts, audit and interrupts.
Updated · 7 min read
What does human-in-the-loop mean?
Human-in-the-loop (HITL) is a control pattern where a person must take part in a decision before an AI system acts on it. For assistants that call tools that read and change real systems, it usually takes the form of an approval gate: the assistant proposes a specific action, execution stops, and a human approves, edits or rejects it.
It differs from human-on-the-loop, where people watch automated actions and can intervene afterwards, and from plain review of AI output, where nothing is executed at all. The gate sits between the model’s intent and the side effect, which is the only point where stopping a mistake is free.
Done well, HITL is selective. Reading a ticket or searching a mailbox runs straight through; sending an email, closing an incident or placing a phone call waits. The rest of this article is about drawing that line and building the pause so it holds under retries, restarts and busy approvers.
Why do assistants need approval gates?
Because models are persuadable and some actions cannot be undone. An assistant reads untrusted content (emails, tickets, web pages) and that content can carry instructions meant to redirect it. Filters and prompts reduce the risk; a gate that requires a person to approve the actual outgoing action removes the worst outcomes.
Accountability is the second reason. When an assistant sends a message to a customer or changes a production record, someone in the organisation must own that decision. An approval ties the action to a named person, with their note, at a known time.
Regulation is moving the same way. Article 14 of the EU AI Act requires high-risk AI systems to be designed so that natural persons can effectively oversee them, including the ability to override or reverse their output and to interrupt them so they halt in a safe state. Not every business assistant is high-risk under the Act, but approval gates are the practical form that oversight takes for tool-using assistants.
Approval modes: never, policy, always
Give every tool a default approval mode, then let policy refine the middle ground. Three modes cover almost every case, and naming them on the tool itself means a new tool is never silently ungoverned.
- Rules can match a tool key or pattern (gmail.*), a risk level, a specific agent, or fields in the tool input such as a recipient address.
- When two rules of equal priority disagree, the stricter one should win: deny beats require approval, which beats allow.
- An individual agent should be able to tighten a tool (require approval where the policy allows) but never loosen it.
| Mode | Meaning | Typical tools |
|---|---|---|
| never | Runs without approval; reserved for read-only tools | Search mail, fetch an incident, query metrics, list changes |
| policy | The tenant’s policy decides per tool, agent and input | Create a draft, add a label, archive, move to trash |
| always | Waits for a human every time by default | Send, reply or forward email; create or update an incident; place a phone call |
Risk levels: read, write, destructive, external_send
Approval modes answer “ask or not”; risk levels explain why. Tagging each tool with one of four levels lets a policy say “everything that leaves the organisation needs approval” once, instead of listing tools one by one.
| Risk level | What it means | Sensible default |
|---|---|---|
| read | Reads data, changes nothing | Allow |
| write | Changes data inside a system you control, usually reversibly | Policy decides; ask unless a rule allows it |
| destructive | Removes or hides data, or is hard to reverse | Require approval |
| external_send | Reaches someone outside: email, message, call | Require approval |
How does an approval pause an assistant run?
In graph-based agent frameworks the pause is an interrupt. LangGraph’s interrupt() suspends the graph inside a node, saves its state through a checkpointer and surfaces a payload to the caller; the graph waits until it is resumed with a Command carrying the human’s answer, which becomes the return value of interrupt(). Its documentation also warns that the node restarts from the beginning on resume, so any code before the interrupt runs again.
That restart is the hard part. Botify handles it this way: its tool node executes exactly one tool call per visit, so a replay can never repeat a sibling call’s side effect, and outgoing actions carry idempotency keys, so an email approved once is sent once.
- The model proposes a tool call; a gateway checks the caller’s identity and roles and evaluates the policy.
- If the policy says approval, a durable approval record is written first, before the graph suspends, so a crash in between still leaves a discoverable request.
- The run is marked as waiting, the pending approval is streamed to the chat, and interrupt() checkpoints the graph state in Postgres.
- An authorised person approves or rejects; the decision is stored independently of the run.
- The run resumes from its checkpoint. On approval the gateway re-reads the decision from storage rather than trusting the caller, then executes the tool.
- On rejection or expiry the model receives a tool result saying the action did not happen, so it can tell the user instead of retrying.
What should the approver see?
Enough to decide in seconds without opening another system, and nothing that should not leave the vault. A good approval card answers: what will happen, to whom, with what content, how risky it is, and until when the request is valid.
- A one-line summary built from the tool and its input, for example “Send mail from Gmail (gmail.send) to finance@… about “Invoice 1042””.
- The tool key and its risk level.
- A preview of the exact payload, with secrets such as access tokens, passwords and API keys redacted and very long text trimmed.
- The expiry time.
- Actions: approve, reject with a note that the model will see, or approve with a corrected payload, which is validated again before execution.
Timeouts: what happens when nobody answers?
Every approval needs an expiry, and expiry must fail closed. A request that nobody answers should end as “not done”, never as “done by default”. In Botify approvals expire after 24 hours unless a policy rule sets its own time, and a background sweep marks overdue requests as expired and resumes each paused run, so it finishes with an honest “the action was not approved in time” instead of hanging forever.
Pick expiry times per action. A reply to a customer may be stale after a few hours; a change to a record can wait a day. Two approvers racing on the same request is another real case: only the first decision should count, which a conditional update on the pending state guarantees.
Audit trails and approval fatigue
Record the whole chain: the request, the policy decision and the rule that produced it, the approval request, who decided, their note, and the result. An append-only, hash-chained log per tenant makes later edits or deletions detectable, which is what an auditor needs to trust it.
The failure mode of HITL is rubber-stamping. If people approve dozens of low-value requests a day, they stop reading them. Keep gates for actions that are irreversible or leave the organisation, let reversible internal writes run under policy once they have proved reliable, and review rejected requests: each one shows where the agent’s judgement or your data needs work.
Frequently asked questions
What is the difference between human-in-the-loop and human-on-the-loop?
In human-in-the-loop, a person must approve an action before it runs. In human-on-the-loop, actions run automatically and people monitor them and intervene afterwards. Agents that send messages or change records usually need the first for those actions.
Does human-in-the-loop slow assistants down?
Only at the gated steps. Reads and reversible work run at full speed; the assistant prepares everything up to the risky action, so the approver decides on a finished proposal rather than doing the work.
What happens to a paused run if the server restarts?
With a durable checkpointer and an approval record written before the pause, nothing is lost. The run resumes from its saved state once the decision arrives, and idempotency keys stop a replay from repeating the action.
Can the approver change what the assistant is about to do?
They should be able to. Approving with a corrected payload, such as a fixed recipient or a shorter sentence, is often faster than rejecting and asking again. The edited input must be validated like any other before it reaches the connected system.
Is human-in-the-loop required by law?
Article 14 of the EU AI Act requires effective human oversight for high-risk AI systems. Many business assistants fall outside that category, but approval gates remain the simplest way to show that a person, not a model, authorised consequential actions.