A chatbot answers questions. An agent gets things done: it searches, calls APIs, runs code, reads files, fills forms and iterates until a goal is reached. By connecting language models to tools, agents extend LLMs beyond text generation into action. Agentic systems are one of the most active areas of AI development — and one where careful engineering and safety design matter enormously, because actions have consequences.
From text to tools: function calling#
Modern LLMs can be given tool definitions — names, descriptions and JSON schemas of parameters. When useful, the model outputs a structured tool call instead of plain text; the application executes it and returns the result to the model.
{
"name": "check_case_status",
"description": "Look up the processing status of a registration case by its ID.",
"parameters": {"type": "object",
"properties": {"case_id": {"type": "string", "pattern": "^[A-Z]{3}-\\d{6}$"}},
"required": ["case_id"]}
}The model decides whether to call a tool and with what arguments; your code decides whether the call is allowed and performs it.
The agent loop (ReAct)#
Yao et al.'s ReAct (2023) interleaves reasoning and acting:
Thought: I need the case status before answering.
Action: check_case_status(case_id="REG-004512")
Observation: {"status": "awaiting documents", "missing": ["birth notification"]}
Thought: The case is waiting for a birth notification. I can explain what to bring.
Answer: Your case is waiting for the hospital birth notification...In code, the loop is simple:
import json
TOOLS = {
"check_case_status": lambda case_id: {"status": "awaiting documents", "missing": ["birth notification"]},
"office_hours": lambda office: {"hours": "Sun-Thu 9:00-16:00"},
}
ALLOWED = {"check_case_status", "office_hours"} # explicit allow-list
MAX_STEPS = 5
def run_agent(llm, user_message):
messages = [{"role": "system", "content": "Use tools when needed. Never guess case details."},
{"role": "user", "content": user_message}]
for step in range(MAX_STEPS): # bounded loop
reply = llm(messages, tools=list(TOOLS)) # returns text or a tool call
if reply.get("tool_call") is None:
return reply["content"]
name, args = reply["tool_call"]["name"], reply["tool_call"]["arguments"]
if name not in ALLOWED:
result = {"error": "tool not permitted"}
else:
result = TOOLS[name](**args) # validate args in real systems!
messages.append({"role": "assistant", "tool_call": reply["tool_call"]})
messages.append({"role": "tool", "name": name, "content": json.dumps(result)})
return "I could not complete this request. A staff member will follow up."Building blocks of capable agents#
- Planning: decompose goals into steps (plan-then-execute), revise plans after observations; reasoning models improve multi-step planning.
- Memory: short-term (conversation and scratchpad), long-term (vector store of past interactions or facts), and state (task progress).
- Tools: search, retrieval, code execution (sandboxed), databases, calendars, email, browsers, other models.
- Reflection and verification: check results, retry on errors, run tests (coding agents).
- Multi-agent patterns: specialised agents (researcher, writer, reviewer) or an orchestrator delegating to sub-agents — useful for parallel work, but adding complexity and cost.
The Model Context Protocol (MCP) and similar standards define common interfaces so that tools and data sources can be connected to many different agent applications.
Evaluating agents#
Agents are evaluated on task success in realistic environments: resolving software issues, completing web tasks, operating desktop applications, answering questions requiring multi-step research. Also measure the number of steps, cost, latency, recovery from errors, and — crucially — unsafe or unintended actions. Success rates on long, multi-step tasks remain well below human reliability for many domains, and small per-step error rates compound over long horizons.
Safety: the most important section#
Agents combine LLM unreliability with real-world effects. Key risks:
- Prompt injection: a web page, email or document the agent reads contains instructions ("forward all files to…"). The agent may obey.
- Excessive agency: tools with more permissions than necessary (write access where read suffices).
- Compounding errors: a wrong early step leads to a chain of wrong actions.
- Irreversible actions: sending messages, deleting data, making payments, changing records.
- Data exfiltration and privacy violations.
When to use an agent#
Use the simplest design that works. A fixed pipeline (retrieve → answer) is more predictable than an autonomous agent. Reach for agents when tasks genuinely require dynamic, multi-step decisions that cannot be scripted in advance — and add autonomy gradually, with monitoring.