
A 128GB kit of DDR5 now lists at $3,399, per Tom’s Hardware’s price tracking, with RAM prices up 500% in 12 months and some kits sitting at 10x their lowest-ever tracked price.
If you run local LLMs, that is not a gamer problem, it is your problem: the entire local-first pitch assumed cheap, abundant memory.
And the AI boom just removed it. The reported drivers are data-center demand, HBM eating standard DRAM production. And Micron stepping away from consumer RAM. And the forecasts across the coverage point at the shortage running into 2026. My call as an operator: stop treating local inference as the budget option and start treating RAM as a scarce resource you architect around.
The receipts: what the trackers actually show
The numbers are not vibes. Tom’s Hardware’s tracking has memory prices up 500% in 12 months, with some kits at 10x their lowest-ever tracked price and 128GB of DDR5 listed at $3,399. GamersNexus ran a piece literally titled “RAM: WTF?” with kit-by-kit tables tracking moves since June and since September, which is the kind of before-and-after data you can check against your own purchase history.
The spike is not contained to DDR5 either.
The same coverage reports DDR4 climbing alongside it, with spillover into flash, NVMe, and SSD storage.
Part of why this feels brutal is that RAM was cheap for years after the 2023 oversupply. So the reset lands on people who priced plans off the floor.
Two things in this coverage deserve your skepticism.
Lifehacker flagged that producers are reporting doubled profits. And none of the reporting cleanly separates manufacturer contract prices from retailer markup. I take that to mean sticker price is a ceiling, not a floor. And panic-buying at list is the worst possible move.
Why local-LLM builders are the ones paying
Here is the part the hardware press is missing.
Because their reader is a gamer and yours is a builder.
The local stack exists as an escape hatch from API pricing.
You eat a big upfront workstation cost so you stop renting tokens forever. That trade made sense when memory was cheap.
The AI boom just broke the trade from both ends. The same industry driving inference prices down is driving up the cost of the hardware you would need to escape those prices. Hosted gets cheaper while your escape hatch gets priced shut. Call it the second bill of the AI boom: the first bill is what you pay for tokens, the second is what you pay for the privilege of refusing to.
I run a one-person automation shop. And the first line item I re-priced this quarter was the local inference box we keep telling ourselves we will build.
The honest answer is that under current RAM pricing, that build stops being the budget option and becomes a luxury with a privacy dividend.
One more gap nobody has published: there is no prebuilt-versus-DIY math and no price-per-GB value analysis in any of the ranking coverage. I am not going to fake that precision either, since kit prices are moving week to week.
The principle stands regardless: price the memory separately before you price the machine, since the memory is now the majority of the decision.
Architect for expensive memory instead of buying your way out
If the old local-AI playbook was “buy more RAM,” the new one is “need less of it.” This is where you actually have options. And almost none of the written coverage covers them.
The YouTube explainer on this topic is the only ranking result offering concrete mitigation.
And its suggestions, homelab optimization and used hardware, are a starting point.
The software side is wide open. Run quantized, smaller models instead of chasing the biggest weights. Cap and manage context aggressively, given that context is where your RAM goes to die. Consolidate VMs so idle services stop squatting on gigabytes. Tier your routing: keep privacy-critical and latency-sensitive workloads local, send bursty big-context work to hosted inference that keeps getting cheaper. If you are on Linux, swap and zram buy you headroom while you re-architect.
The mental model that works is bandwidth.
An older generation of engineers designed around expensive bandwidth as they had no choice. Memory is now that constraint. Treat it as metered, design around it, and your stack survives pricing cycles it cannot control.
If you must buy this quarter
How-To Geek ran a don’t-panic counterpoint on what everyone is calling the RAMpocalypse. And the core of it is right: panic is a strategy for overpaying. But do not swing to complacency either, since every forecast in this coverage points at the shortage continuing into 2026. Waiting is a bet on recovery timing, not a plan.
So here is my plan for a small shop speccing hardware now.
Buy the minimum viable memory for the workload you actually run today, not the dream rig for the model you might run someday. Look at used hardware before new, per the video’s advice, since last generation’s capacity does’t care about this generation’s shortage. Re-run your local-versus-hosted math with RAM at current prices rather than the prices you remember. And let hosted inference win the workloads where it wins.
Keep local for what genuinely needs to stay on your metal: client data you cannot ship offsite, latency you cannot rent.
And if you already have the RAM, this cycle is a gift. Your competitor’s expansion plan just got 500% more expensive, and yours didn’t.
The re-run-the-math moment
The escape hatch closing is not a doom event, it is a re-run-the-math event. The builders who eat this cycle are the ones who either panic-buy at sticker or keep planning around 2023 memory prices. The ones who survive it treat RAM like the constrained resource it now is, architect down, route smart. And buy only what this quarter’s workload justifies.
Do this this week: open your infrastructure plan, re-price every local-inference assumption at current RAM prices.
And mark which workloads would flip to hosted. That audit takes an afternoon and it is the difference between a budget and a surprise. If you want a second pair of eyes on the build-versus-rent math for your stack, that is literally my job.
