In the first lecture we defined AI as the construction of rational agents. Today we make that definition precise. By the end you should be able to take any AI problem — a trading bot, a medical assistant, a drone — and describe it in the vocabulary of agents and environments. This vocabulary is the lingua franca of the entire field.
Agents, percepts and actions#
An agent is anything that perceives its environment through sensors and acts upon it through actuators. A human has eyes and hands; a robot has cameras and motors; a software agent receives keystrokes or network packets and outputs text or API calls.
- A percept is the agent's input at a single instant.
- The percept sequence is the complete history of everything the agent has perceived.
- The agent function maps percept sequences to actions.
- The agent program is the concrete implementation running on physical hardware.
The distinction between function and program matters: the function is a mathematical object (possibly infinite), while the program must be finite and run in bounded time.
Rationality, formally#
A rational agent selects, for every possible percept sequence, the action expected to maximise its performance measure, given the evidence of the percept sequence and whatever built-in knowledge it has.
Here $U$ is the performance (utility), $p_{1:t}$ is the percept history and $K$ is prior knowledge. Note three subtleties:
- Rational is not omniscient. An agent that crosses the street after looking carefully and is hit by a falling satellite was still rational.
- Rationality includes information gathering. Looking before crossing is a rational action because it improves future decisions.
- Rationality requires autonomy. An agent relying only on its designer's prior knowledge, never learning, is fragile.
The PEAS description#
To specify a task environment, list its Performance measure, Environment, Actuators and Sensors.
| Agent | Performance | Environment | Actuators | Sensors |
|---|---|---|---|---|
| Automated taxi | Safety, speed, legality, comfort, profit | Roads, traffic, pedestrians, weather | Steering, accelerator, brake, horn, display | Cameras, LiDAR, GPS, speedometer |
| Medical diagnosis assistant | Patient health, cost, lawsuits avoided | Patient, hospital staff | Questions, test orders, diagnoses | Symptoms, test results, patient answers |
| Spam filter | Accuracy, low false positives | Email stream, users | Label as spam / not spam | Email text, headers, metadata |
| Refugee registration chatbot | Correct answers, accessibility, privacy | Users in many languages, case database | Text replies, referrals | Typed or spoken messages |
Properties of environments#
Environments differ along several dimensions, and each dimension determines which algorithms are appropriate.
- Fully vs partially observable. Can the sensors see the complete relevant state? Chess: fully. Poker: partially.
- Single vs multi-agent. Are other agents optimising their own goals? Multi-agent environments may be competitive or cooperative.
- Deterministic vs stochastic. Is the next state completely determined by the current state and action?
- Episodic vs sequential. Is each decision independent (classifying images) or do actions affect the future (driving)?
- Static vs dynamic. Does the world change while the agent deliberates?
- Discrete vs continuous. Are states, time and actions countable?
- Known vs unknown. Does the agent know the rules (the transition model)?
The hardest case — partially observable, multi-agent, stochastic, sequential, dynamic, continuous and unknown — describes much of real life, including driving a taxi.
Agent architectures#
We now study four agent designs of increasing sophistication.
1. Simple reflex agents#
Act only on the current percept using condition–action rules. Fast and simple, but they fail when the environment is partially observable.
2. Model-based reflex agents#
Maintain an internal state that tracks aspects of the world the agent cannot currently see, updated using a model of how the world evolves and how actions affect it.
3. Goal-based agents#
Know what states are desirable and use search or planning to find action sequences that reach them. More flexible: change the goal and behaviour changes without rewriting rules.
4. Utility-based agents#
Replace a binary goal with a utility function that scores how desirable each state is, allowing trade-offs (fast vs safe) and decisions under uncertainty via expected utility.
Learning agents#
Any of these can be made into a learning agent with four components: a performance element that chooses actions, a critic that evaluates outcomes against a standard, a learning element that improves the performance element, and a problem generator that suggests exploratory actions. Reinforcement learning, which we study later, is a precise mathematical realisation of this picture.
class ModelBasedAgent:
def __init__(self, rules, update_state):
self.state = {}
self.last_action = None
self.rules = rules # list of (condition_fn, action)
self.update_state = update_state
def __call__(self, percept):
self.state = self.update_state(self.state, self.last_action, percept)
for condition, action in self.rules:
if condition(self.state):
self.last_action = action
return action
self.last_action = "NoOp"
return "NoOp"