Illustration: massive gleaming GPU fortress looming over a wheezing laptop with a desk fan pointed at it, flat vector, warm reds/oranges
There are two kinds of people in AI now: the GPU rich and the GPU poor. The GPU rich train frontier models on clusters the size of small towns and argue about power substations. The GPU poor refresh a rate-limit dashboard, do arithmetic about tokens per dollar, and have strong opinions about quantization. This is the most honest class divide in tech β nobody pretends it doesn't exist, and everyone knows exactly which side they're on.
What "GPU Rich" Actually Looks Like
The GPU rich don't buy GPUs; they buy power plants adjacent to GPUs. xAI's Colossus cluster in Memphis scaled past 100,000 H100s β a quarter-billion dollars of silicon in one building, drinking a city's worth of electricity. Meta, Google, OpenAI, and Anthropic operate at similar scale. When these companies say "we're compute-constrained," they mean they have merely enormous amounts of compute, and the stock market treats this as a tragedy.
Then there's the landlord class: NVIDIA, which briefly became the most valuable company on Earth by selling shovels in a gold rush it helped start, and the neoclouds β CoreWeave went from crypto mining to a multi-billion-dollar AI cloud via the ancient business strategy of buying GPUs before everyone else realized they needed them. The GPU rich don't worry about per-token pricing. They worry about lead times and whether the local grid can handle their ambitions.
The GPU Poor Economy
The GPU poor β which is to say, nearly every startup, researcher, indie hacker, and curious weirdo β live in a different universe:
- Rent by the hour. An H100 rents for a few dollars an hour on spot markets. Your training run costs real money every minute it exists, which concentrates the mind wonderfully.
- Inference APIs. Why own the cow when you can rent the milk by the token? Providers sell frontier-model inference at prices the poor can afford, which is the entire business model of half of Y Combinator.
- The people's cluster. A MacBook running a local model, a dusty 3090 in a gaming PC, a free notebook that disconnects if you look away. The GPU poor have turned "will it run on my laptop" into a research discipline.
- Quantization and LoRA. The great equalizers. QLoRA let researchers fine-tune 65-billion-parameter models on a single GPU β the technical equivalent of fitting an elephant into a Mini Cooper and driving it to work.
The phrase "GPU poor" wasn't coined by a journalist β it emerged from ML researchers themselves, half joke, half cri de coeur. It's the rare piece of industry slang that's also a precise economic descriptor.
The Divide Shapes What Gets Built
Here's the part nobody puts in the pitch deck: compute inequality determines what kind of AI work is possible for whom. Only the rich can pretrain. The poor do everything downstream β evals, agents, fine-tunes, distillation, the clever hacks that squeeze more from less. An uncomfortable amount of "open-source AI progress" is really "GPU-poor researchers finding brilliant ways to cope."
This is why open weights matter so much. Every major open release is a wealth transfer from the GPU rich to the GPU poor β billions of dollars of training compute, handed out as a download. The rich companies know this and do it anyway: partly goodwill, partly ecosystem capture, partly because their researchers are romantics. The GPU poor should be grateful and also slightly suspicious, which is the correct attitude toward all gifts from billionaires.
The talent market runs on the same fault line. The GPU rich hire by buying entire labs β acqui-hires where the real asset is the researchers plus the cluster access that comes with them. The GPU poor compete by publishing: a clever efficiency trick from a grad student with one A100 can land a job offer from a frontier lab, which then hands them the cluster they never had. It's the purest meritocracy in tech and the most lopsided one β your ideas have to be ten times better to compensate for having a thousand times less compute.
Meanwhile, the flood of machine-made content drowning the internet is, not coincidentally, a GPU-poor phenomenon at scale: inference got cheap enough that generating a million mediocre articles costs less than hiring one writer. And when the models everyone depends on can be jailbroken with clever wording, the poor inherit the rich's security headaches too β without the rich's security teams.
How to Play It If You're Poor
Rent, don't buy
Unless your burn rate has nine digits, you will never win the capex game. Rent spot instances, use inference APIs, and spend your money on the things compute can't buy: data, distribution, and taste.
Distill everything
The frontier labs spend nine figures training a model; you spend three figures distilling it into something that runs on a laptop. Distillation is the have-nots' judo β use the giant's strength against it. (The giants have opinions about this, which tells you it works.)
Bet on efficiency, not scale
Every point of algorithmic efficiency is a tax cut for the poor. Quantization, speculative decoding, better architectures β the history of the compute-poor is the history of doing more with less, and the "less" keeps getting more capable every year.
Remember what the moat actually is
Compute is rented. Weights leak. The durable advantages are proprietary data, distribution, and being the thing users actually love. Nobody ever won by having the second-biggest cluster. The GPU poor have one structural advantage the rich can never buy: they have to be clever.
Enjoyed this? Get the next one first.
The Spicy Dispatch: one email a week, zero slop, unsubscribe anytime.
One email a week. Zero slop. Unsubscribe anytime.