AI is often imagined as weightless software, but it runs on physical infrastructure: chips manufactured with energy- and water-intensive processes, data centres drawing large amounts of electricity, and cooling systems that consume water. As model sizes and usage grow, so do these environmental costs. Understanding them is part of responsible engineering — and efficiency is also good engineering, reducing cost and widening access.
Where the costs come from#
- Training: large, one-off (but repeated across experiments) computations. Strubell, Ganesh and McCallum (2019) drew attention to the energy cost of training NLP models, including extensive hyperparameter and architecture searches.
- Inference: each query is cheap, but billions of queries add up — for widely used models, cumulative inference energy can exceed training energy.
- Embodied emissions: manufacturing GPUs, servers and data-centre buildings.
- Water: data centres use water for cooling, directly and indirectly through electricity generation; studies (e.g. Li et al., 2023) have estimated significant water use associated with large-model training and use.
- Experimentation overhead: failed runs, tuning and ablations often use more compute than the final training run.
Estimating the footprint#
A standard estimate of operational emissions:
- PUE (Power Usage Effectiveness) accounts for data-centre overhead such as cooling (1.1 for efficient hyperscale facilities; higher for typical server rooms).
- Carbon intensity of electricity varies enormously by region and time — from a few tens of grams per kWh on hydro- or nuclear-heavy grids to 700+ g/kWh on coal-heavy grids.
def training_emissions(gpu_hours, gpu_power_kw=0.4, pue=1.2, carbon_kg_per_kwh=0.5):
energy_kwh = gpu_hours * gpu_power_kw * pue
return energy_kwh, energy_kwh * carbon_kg_per_kwh
for name, hours, intensity in [("fine-tune small model", 8, 0.5),
("train mid-size model, coal-heavy grid", 5_000, 0.7),
("same, low-carbon grid", 5_000, 0.05)]:
kwh, kg = training_emissions(hours, carbon_kg_per_kwh=intensity)
print(f"{name:<40} {kwh:>9,.0f} kWh {kg:>9,.0f} kg CO2e")The location and timing of compute can change emissions by an order of magnitude. Tools such as CodeCarbon and experiment-impact trackers measure energy during training automatically.
What drives costs up#
- Scale of models and data (training compute ≈ $6ND$).
- Inefficient hardware utilisation (idle GPUs waiting for data).
- Large hyperparameter searches and repeated runs.
- Serving oversized models for tasks that small ones could handle.
- Always-on infrastructure and forgotten cloud instances.
Towards "Green AI"#
Schwartz et al. (2020) proposed "Green AI": reporting efficiency (e.g. FLOPs or energy) alongside accuracy, and valuing efficiency improvements as research contributions — not only accuracy at any cost.
Practical steps:
- Right-size models: use the smallest model that meets requirements; try classical ML before deep learning for tabular data.
- Reuse: fine-tune pretrained models (and use parameter-efficient methods like LoRA) instead of training from scratch.
- Efficient training: mixed precision, good data loading, early stopping, smarter hyperparameter search (Bayesian optimisation, Hyperband) instead of huge grids.
- Efficient inference: quantisation, distillation, batching, caching, routing easy queries to small models.
- Choose low-carbon regions and times where possible; prefer efficient data centres.
- Measure and report energy and emissions in papers, theses and project reports.
- Hardware lifecycle: extend hardware life, avoid unnecessary upgrades, responsible recycling.
A balanced view#
Estimates of AI's total energy share vary widely and change quickly as usage and efficiency evolve. Avoid both dismissal ("it's negligible") and exaggeration; instead, measure your own systems, report transparently, and make efficient choices that cost little in quality.