Few topics generate more headlines โ and more confusion โ than "Artificial General Intelligence". As future scientists and engineers, you will be asked about it by journalists, policymakers and your own families. Today we build a careful vocabulary so you can discuss it with precision rather than hype.
Narrow AI#
Narrow (or weak) AI refers to systems designed or trained for a specific task or a bounded family of tasks: recognising faces, translating text, recommending videos, playing Go. Narrow systems can be superhuman within their domain โ no human beats AlphaZero at chess โ yet useless outside it. AlphaZero cannot tell you what a chess piece is made of.
General intelligence#
Artificial General Intelligence (AGI) usually refers to a system that can perform a wide range of cognitive tasks at or above human level, adapt to new tasks efficiently, and transfer knowledge across domains. The key word is generality โ not peak performance on any single task.
Several definitions coexist:
- Human-comparison definitions: AGI can do most economically valuable cognitive work that humans can do.
- Capability-breadth definitions: performance across a broad battery of tasks.
- Learning-efficiency definitions: Franรงois Chollet argues that intelligence is skill-acquisition efficiency โ how quickly a system learns new skills from little data, relative to its priors and experience. His ARC (Abstraction and Reasoning Corpus) benchmark tests this with novel visual puzzles.
- Levels frameworks: some researchers propose graded levels (emerging, competent, expert, virtuoso, superhuman) crossed with breadth (narrow vs general), treating AGI as a spectrum rather than a single threshold.
Where current systems stand#
Large language and multimodal models are the most general AI systems built so far: one model writes code, summarises law, explains physics and describes images. At the same time, careful evaluations reveal characteristic weaknesses:
- sensitivity to phrasing and irrelevant details;
- difficulty with problems requiring long chains of reliable reasoning or genuinely novel abstractions;
- factual errors stated confidently (hallucination);
- limited ability to learn continually from experience after training.
The honest scientific position is that generality has increased dramatically, measurement is hard, and experts disagree about how far current approaches will go.
Measuring generality#
A good evaluation of general capability should be:
- Broad โ covering many domains and skill types.
- Novel โ tasks the system could not have memorised from training data (data contamination is a serious problem when models are trained on much of the internet).
- Efficiency-aware โ accounting for how much data and compute the system needed.
- Robust โ performance should survive rephrasing and adversarial variation.
# A tiny illustration of contamination checking: n-gram overlap between
# a benchmark question and a training corpus.
def ngrams(text, n=8):
words = text.lower().split()
return {" ".join(words[i:i + n]) for i in range(len(words) - n + 1)}
def overlap_ratio(test_item, corpus_ngrams, n=8):
grams = ngrams(test_item, n)
return len(grams & corpus_ngrams) / max(1, len(grams))Superintelligence#
Superintelligence refers to intellect that greatly exceeds human cognitive performance in virtually all domains. Arguments about it include:
- The intelligence explosion hypothesis (I. J. Good, 1965): a machine that can improve its own design could trigger recursive self-improvement.
- Orthogonality thesis (Bostrom): intelligence and final goals are independent โ a highly capable system need not share human values.
- Instrumental convergence: many goals imply similar sub-goals (acquire resources, avoid being switched off), which could create risks if goals are misspecified.
Sceptics respond that intelligence is not a single scalar, that real-world progress is bottlenecked by experiments, data and physical resources, and that capabilities have historically grown gradually rather than explosively.
A responsible stance#
As a researcher, you can hold two ideas at once:
- Present-day harms โ bias, misinformation, privacy violations, labour disruption โ are real and need solving now.
- Longer-term risks from increasingly capable systems deserve serious technical research in alignment, interpretability and evaluation.
These are not competing priorities; both demand rigorous engineering and good governance. We return to both in the ethics track.