A scientist who does not know the history of their field is condemned to rediscover old ideas and repeat old mistakes. Today we take a guided tour through the history of Artificial Intelligence. Pay attention to a recurring pattern: bold promises, brilliant early results, a collision with reality, and then a new idea that changes the game.
Prehistory: the dream of mechanical reasoning (before 1950)#
The idea that reasoning could be mechanised is ancient. Aristotle formalised syllogisms. In the 17th century Leibniz dreamed of a calculus ratiocinator that would settle arguments by calculation. George Boole (1854) showed that logic could be written as algebra, and Gottlob Frege created predicate logic. In 1936 Alan Turing defined the Turing machine, giving a precise meaning to "computation" itself.
Two more ingredients arrived in the 1940s. Warren McCulloch and Walter Pitts (1943) proposed a mathematical model of a neuron as a threshold logic unit and showed that networks of such units could compute logical functions. Norbert Wiener founded cybernetics, the study of feedback and control. Both ideas โ logic and neurons โ would compete and cooperate for the next eighty years.
Birth of the field (1950โ1956)#
In 1950 Turing published Computing Machinery and Intelligence, asking "Can machines think?" and proposing the imitation game. In 1956 John McCarthy, Marvin Minsky, Claude Shannon and Nathaniel Rochester organised the Dartmouth Summer Research Project, where the phrase "Artificial Intelligence" was adopted. Their proposal famously conjectured that "every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it."
The golden years (1956โ1974)#
Early successes were dazzling:
- Logic Theorist (Newell and Simon, 1956) proved theorems from Principia Mathematica.
- General Problem Solver tried to capture human-like meansโends reasoning.
- Arthur Samuel's checkers program (1959) learned to beat its creator โ and Samuel coined the term machine learning.
- Frank Rosenblatt's Perceptron (1958) learned to classify simple patterns in hardware.
- ELIZA (Weizenbaum, 1966) simulated a psychotherapist with simple pattern matching and surprised people with how human it felt.
- Shakey the robot combined perception, planning and action.
Researchers predicted human-level AI within a generation.
The first AI winter (1974โ1980)#
Reality intervened. Programs that worked on toy problems failed to scale due to combinatorial explosion. Machine translation performed poorly. In 1969 Minsky and Papert's book Perceptrons proved that a single-layer perceptron cannot represent functions such as XOR, which dampened neural network research. The 1973 Lighthill Report in the UK criticised the field, and funding was cut on both sides of the Atlantic.
Expert systems boom and bust (1980โ1993)#
AI revived commercially through expert systems โ programs encoding the ifโthen rules of human specialists. MYCIN diagnosed blood infections; XCON configured computer orders and saved its company millions. Japan launched the Fifth Generation Computer project. But expert systems were brittle, expensive to maintain and could not learn. When specialised Lisp machine hardware collapsed commercially in the late 1980s, a second AI winter followed.
Meanwhile, something important happened quietly: in 1986 Rumelhart, Hinton and Williams popularised backpropagation for training multi-layer networks, solving the problem Minsky and Papert had highlighted.
The statistical turn (1990sโ2000s)#
The field matured by embracing probability and statistics. Judea Pearl's Bayesian networks gave a principled way to reason under uncertainty. Support Vector Machines, boosting and random forests achieved excellent results with strong theory. IBM Deep Blue defeated Garry Kasparov in 1997 using massive search, not learning. Yann LeCun's convolutional networks read handwritten cheques. Researchers increasingly measured progress on shared benchmark datasets.
The deep learning revolution (2006โ2017)#
Three forces converged: large datasets (ImageNet, 2009), GPU computing, and better training techniques (ReLU activations, dropout, good initialisation). In 2012, AlexNet won the ImageNet challenge by a huge margin, and the field changed almost overnight. Milestones followed quickly:
| Year | Milestone |
|---|---|
| 2013 | Word2Vec learns word meanings as vectors; DQN learns Atari games from pixels |
| 2014 | Generative Adversarial Networks; sequence-to-sequence translation |
| 2015 | ResNet trains networks with 150+ layers |
| 2016 | AlphaGo defeats Lee Sedol at Go |
| 2017 | "Attention Is All You Need" introduces the Transformer |
The era of foundation models (2018โpresent)#
The Transformer enabled models pretrained on enormous corpora: BERT (2018) for understanding and the GPT series for generation. Scaling laws showed that performance improves predictably with more data, parameters and compute. Diffusion models transformed image generation. AlphaFold 2 (2020) effectively solved protein structure prediction for many proteins. Conversational assistants built on large language models reached hundreds of millions of users. Today the research frontier includes reasoning, multimodality, efficiency, agents that use tools, and โ crucially โ safety and alignment.
Lessons for a young researcher#
- Beware of over-promising. Both AI winters followed inflated expectations.
- Old ideas return. Neural networks were declared dead twice. Backpropagation, attention and reinforcement learning all have deep roots.
- Benchmarks drive progress โ and can mislead. Always ask what a benchmark does not measure.
- Compute and data matter as much as cleverness.