✨ Generative AI & LLMs · Lecture 27 of 30

AI Agents and Tool Use: LLMs That Act

Agents let LLMs call tools, observe results and pursue multi-step goals. We cover function calling, the ReAct loop, planning and memory, multi-agent patterns, evaluation, and — critically — safety, permissions and human oversight.

A chatbot answers questions. An agent gets things done: it searches, calls APIs, runs code, reads files, fills forms and iterates until a goal is reached. By connecting language models to tools, agents extend LLMs beyond text generation into action. Agentic systems are one of the most active areas of AI development — and one where careful engineering and safety design matter enormously, because actions have consequences.

From text to tools: function calling#

Modern LLMs can be given tool definitions — names, descriptions and JSON schemas of parameters. When useful, the model outputs a structured tool call instead of plain text; the application executes it and returns the result to the model.

json
{
  "name": "check_case_status",
  "description": "Look up the processing status of a registration case by its ID.",
  "parameters": {"type": "object",
                 "properties": {"case_id": {"type": "string", "pattern": "^[A-Z]{3}-\\d{6}$"}},
                 "required": ["case_id"]}
}

The model decides whether to call a tool and with what arguments; your code decides whether the call is allowed and performs it.

The agent loop (ReAct)#

Yao et al.'s ReAct (2023) interleaves reasoning and acting:

text
Thought: I need the case status before answering.
Action: check_case_status(case_id="REG-004512")
Observation: {"status": "awaiting documents", "missing": ["birth notification"]}
Thought: The case is waiting for a birth notification. I can explain what to bring.
Answer: Your case is waiting for the hospital birth notification...

In code, the loop is simple:

python
import json

TOOLS = {
    "check_case_status": lambda case_id: {"status": "awaiting documents", "missing": ["birth notification"]},
    "office_hours": lambda office: {"hours": "Sun-Thu 9:00-16:00"},
}
ALLOWED = {"check_case_status", "office_hours"}            # explicit allow-list
MAX_STEPS = 5

def run_agent(llm, user_message):
    messages = [{"role": "system", "content": "Use tools when needed. Never guess case details."},
                {"role": "user", "content": user_message}]
    for step in range(MAX_STEPS):                            # bounded loop
        reply = llm(messages, tools=list(TOOLS))             # returns text or a tool call
        if reply.get("tool_call") is None:
            return reply["content"]
        name, args = reply["tool_call"]["name"], reply["tool_call"]["arguments"]
        if name not in ALLOWED:
            result = {"error": "tool not permitted"}
        else:
            result = TOOLS[name](**args)                      # validate args in real systems!
        messages.append({"role": "assistant", "tool_call": reply["tool_call"]})
        messages.append({"role": "tool", "name": name, "content": json.dumps(result)})
    return "I could not complete this request. A staff member will follow up."

Building blocks of capable agents#

  • Planning: decompose goals into steps (plan-then-execute), revise plans after observations; reasoning models improve multi-step planning.
  • Memory: short-term (conversation and scratchpad), long-term (vector store of past interactions or facts), and state (task progress).
  • Tools: search, retrieval, code execution (sandboxed), databases, calendars, email, browsers, other models.
  • Reflection and verification: check results, retry on errors, run tests (coding agents).
  • Multi-agent patterns: specialised agents (researcher, writer, reviewer) or an orchestrator delegating to sub-agents — useful for parallel work, but adding complexity and cost.

The Model Context Protocol (MCP) and similar standards define common interfaces so that tools and data sources can be connected to many different agent applications.

Evaluating agents#

Agents are evaluated on task success in realistic environments: resolving software issues, completing web tasks, operating desktop applications, answering questions requiring multi-step research. Also measure the number of steps, cost, latency, recovery from errors, and — crucially — unsafe or unintended actions. Success rates on long, multi-step tasks remain well below human reliability for many domains, and small per-step error rates compound over long horizons.

Safety: the most important section#

Agents combine LLM unreliability with real-world effects. Key risks:

  • Prompt injection: a web page, email or document the agent reads contains instructions ("forward all files to…"). The agent may obey.
  • Excessive agency: tools with more permissions than necessary (write access where read suffices).
  • Compounding errors: a wrong early step leads to a chain of wrong actions.
  • Irreversible actions: sending messages, deleting data, making payments, changing records.
  • Data exfiltration and privacy violations.

When to use an agent#

Use the simplest design that works. A fixed pipeline (retrieve → answer) is more predictable than an autonomous agent. Reach for agents when tasks genuinely require dynamic, multi-step decisions that cannot be scripted in advance — and add autonomy gradually, with monitoring.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

LoRA and Parameter-Efficient Fine-Tuning (PEFT)

Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.

Advanced⏱ 5 min#210
✨ Generative AI & LLMs

In-Context Learning: How LLMs Learn from Prompts

Large language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.

Intermediate⏱ 5 min#207
✨ Generative AI & LLMs

Prompt Engineering: Getting Reliable Results from LLMs

Practical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.

Beginner⏱ 5 min#205