
Image: Olkeri
By Olkeri.space
AI Chips Explained: Why GPUs Run the World's Artificial Intelligence
Why graphics chips became the engine of artificial intelligence, how GPUs, TPUs and NPUs differ, and why memory bandwidth decides real AI performance.
Read this story in: Français · Deutsch · Español
Artificial intelligence runs on a narrow, expensive and strategically contested class of hardware. Understanding why explains a great deal about the AI industry: why one chipmaker became one of the world's most valuable companies, why governments treat chip exports as national security, and why compute capacity is the real constraint on AI progress.
Why graphics chips, of all things:
Training a neural network is, mathematically, an enormous pile of matrix multiplication. The same simple arithmetic operation is repeated billions of times on different numbers.
Central processing units, CPUs, are built for the opposite problem. A CPU has a handful of powerful cores optimised to run complicated, unpredictable instruction sequences quickly, one after another. That is ideal for operating systems and most software, and poor for doing one simple operation billions of times.
Graphics processing units were designed to calculate the colour of millions of pixels simultaneously, so they have thousands of simpler cores built for doing the same thing in parallel. That architecture, developed for video games, turned out to be almost exactly what neural networks need.
The decisive factor was software. Nvidia released CUDA in 2007, letting developers program GPUs for general computation. By the time deep learning took off around 2012, an entire ecosystem of tools, libraries and trained engineers already existed. Competitors have built comparable silicon since; matching that software ecosystem has proven far harder.
The alphabet soup: GPU, TPU, NPU and ASIC:
A GPU is the general-purpose workhorse of AI. Flexible enough to train any model architecture, powerful enough for the largest workloads, and available in cloud data centres worldwide.
A TPU, or tensor processing unit, is Google's custom chip built specifically for neural network operations. Specialising narrows flexibility but improves efficiency for the workloads it targets. Several large technology companies now design their own equivalents, partly for performance and partly to reduce dependence on a single supplier.
An NPU, or neural processing unit, is a small, low-power accelerator built into phones, laptops and cameras. It runs modest models locally so that photo processing, voice recognition and translation work without sending data to a server. That improves latency, privacy and offline capability.
ASIC is the general term for any chip designed for one specific task. TPUs and NPUs are both examples. The trade-off is always the same: the narrower the purpose, the better the efficiency, and the greater the risk that the workload changes and the silicon no longer fits.
Training and inference are different problems:
Training builds the model. It is a one-off, enormous computation requiring thousands of chips working together for weeks, consuming megawatts of power. The cost of training a frontier model runs into the tens or hundreds of millions of dollars.
Inference is using the trained model to answer a request. Each individual inference is small, but a popular product performs billions of them. Over a successful model's lifetime, the total money and energy spent on inference typically exceeds what was spent on training.
The two workloads reward different hardware. Training rewards raw throughput and fast interconnects between many chips. Inference rewards low latency, memory capacity and cost per query. Increasingly, chips are designed for one or the other rather than both.
Memory bandwidth is the real bottleneck:
Chip marketing emphasises raw computational throughput, but in practice the limiting factor is usually memory bandwidth: how fast data moves between memory and the processing cores.
Modern accelerators can perform arithmetic far faster than memory can supply numbers to work on. Cores sit idle waiting for data. This is why high bandwidth memory, or HBM, matters so much. HBM stacks memory chips vertically and connects them with very wide, very fast paths, and its limited supply has repeatedly constrained the entire AI hardware market. A shortage of memory, not of processors, has been the binding constraint in several recent periods.
Chips do not work alone:
A frontier model is trained across thousands of chips simultaneously, which means they must constantly exchange results. If the network connecting them is slow, expensive accelerators wait.
This is why networking has become central to AI infrastructure, and why complete systems, entire racks with chips, memory, networking and cooling designed together, are now sold as single products. Performance is a property of the system, not of any one component.
Power and cooling became the constraint:
The practical limit on AI expansion is increasingly electricity. A large AI data centre can draw as much power as a small city, and the industry's growth has collided with the pace at which grids can be expanded and connections approved.
Heat follows power. Dense AI racks have pushed the industry from air cooling toward liquid cooling, which changes how data centres are built. Operators now choose sites for power availability and water access as much as for network proximity, which is why nuclear, geothermal and dedicated generation deals have become part of AI strategy.
Why chips became geopolitics:
The supply chain has extreme chokepoints. The most advanced chips are manufactured by a small number of foundries, overwhelmingly in Taiwan and South Korea. The machines that make them, particularly extreme ultraviolet lithography systems, come from essentially one company in the Netherlands. Much of the design software is American.
That concentration turned semiconductors into an instrument of policy. Export controls restrict which chips can be sold to which countries; subsidy programmes attempt to bring manufacturing onshore; and a shortage or a shipping disruption becomes an economic event.
What it means practically:
Compute is the scarce input in artificial intelligence. Access to it determines who can train frontier models, how quickly products improve, and what they cost to run.
For most organisations the implication is simple: you will rent this capacity rather than own it, your costs will be driven by inference volume rather than training, and the efficiency of the model you choose will matter more to your budget than any headline benchmark.