Skip to content
Skip to content
● ENGINE LIVESign inCreate an account

wubbery://memory — the decision layer above the memory hierarchy

HBM4 is extraordinary. It is also not enough on its own — and that is not a criticism, it is arithmetic.

The memory industry is having its best decade in forty years. Micron is shipping HBM4 at 2.8 TB/s and is sold out for the year. NVIDIA has standardised a three-tier context hierarchy. Researchers are demonstrating six. We want every one of them to win, because none of it decides which byte lives where — and that decision is now worth more than the bandwidth.

“WUBBERY wins, and we’ll never stop growing until the wheel we just reinvented is attached to the vehicle that is driving innovation for the entire world.”
Tim Harkin, founder
01

What the hardware industry just achieved

We are going to spend this page explaining a gap, so let us be precise about how much of it has already been closed by other people, and how impressive that work is. Every figure in this section is theirs, published by them, and linked.

That is a genuine generational step, delivered against real physics by people who had to solve stacking, thermals and yield to get there. None of it is marketing. We have built our memory work on the assumption that all of it ships and all of it works.

02

And the gap it does not close

320 GB
KV cache for one million-token context on a 70B model
80 GB
HBM on the accelerator it runs on
The shortfall. No plausible stack height closes it

This is why the entire industry moved to tiering, and tiering is the point. A hierarchy of HBM, DRAM, CXL and NVMe is not a capability — it is a decision surface. Every byte placed on the wrong tier is either latency a user feels or capacity bought and wasted, and there are billions of those decisions per second.

Nobody sells that decision, because it is not silicon. It does not fabricate, it does not stack, and it cannot be bought by the wafer. It is the layer above, and it is the one we have spent our time on.

03

We are not competing with memory. We make it pay.

There is a version of this business that positions against the hardware vendors. It would be both dishonest and stupid. Every dollar they spend making memory faster makes our product more valuable and none of it makes us less necessary. We would like them to win. We are part of the reason winning pays.

The clearest case is our flagship module, which exists to decline a purchase. memory.bandwidthBound tells a customer that their workload is limited by memory and that additional compute will return nothing. That is not a contrarian position — it is exactly what the memory industry has been arguing for years. The difference is that we prove it for one specific workload, with arithmetic the customer can reproduce in the room, in about four microseconds.

A memory vendor telling you memory is the bottleneck is a sales pitch. A vendor-neutral measurement telling you the same thing is evidence. We are happy to be the evidence.

And when the measurement finds a limit, we do not stop there — we go and remove it. Every hard answer this engine gives is paired with the module that changes it, both shipped, both measured. That is the next section, and it is the whole reason this engine exists.

04

Seventeen memory modules, every figure measured

Nine of these refuse, and six exist purely to remove a refusal. That is deliberate: a refusal is hard to fake, easy to demonstrate, and worth more than a percentage to anyone signing off a purchase. Each is benchmarked against a case that passes the check the industry normally runs — a guard that only blocks obvious nonsense is measuring nothing.

memory.bandwidthBound2/2 blocked · +100%

Refuse the compute purchase that returns nothing

Arithmetic intensity of 2 FLOP/byte against an 89.3 ridge point. Autoregressive decode reaches 2.2% of installed compute. Doubling compute returns 0.0%; doubling bandwidth returns 100%.

memory.kvAdmit2/2 blocked

Admit on the final footprint, not the prompt

A request that fits at admission and not at completion is admitted by every naive controller, and dies near the end of a generation the customer has already paid for and waited through.

memory.tierPlace3/3 blocked

Place by reuse distance, price every miss

A tier is not free capacity, it is capacity with a latency attached. Blocks are ordered by when they are next needed, and a placement is refused if a block due inside the deadline cannot be delivered in time.

memory.rehydrateBudget2/2 blocked · 38.6× gap

The datasheet belongs to a single reader

A tier that meets its time-to-first-token target at concurrency one misses it by 38.6× at forty streams, because the link is shared. This is why tiering benchmarks reproduce so badly.

memory.migrationHysteresis3/3 blocked

Refuse the promotion that never amortises

Promotion costs bandwidth on both tiers. Get the decision wrong and the page thrashes between them, until migration traffic exceeds the traffic it was meant to accelerate.

memory.strandedDram2/2 blocked · 90.6% excluded

Procurement error is not a pooling benefit

Memory above observed peak is a buying mistake you fix by buying less. Only the band between mean and peak can ever be pooled. Quoting the total counts the mistake twice.

memory.cxlEconomics2/2 blocked

Answer the objection with arithmetic

Pooling pays by statistical multiplexing, so the deciding term is peak correlation. On a fleet running one workload every host peaks together and the benefit vanishes. Returns the break-even correlation for your fleet.

memory.hbmHealthDrain3/3 blocked

Drain before it takes the whole run down

Weighs the capacity cost of draining against expected rank-hours lost. The same degrading device is drained before a 54-day run and left in service for a two-hour one.

memory.bandwidthFairShare2/2 blocked

Capacity fairness is not bandwidth fairness

A tenant inside its memory quota can saturate the controller and inflate every neighbour's latency. Every tenant is compliant, so a capacity audit finds nothing.

memory.peakDecorrelate2/2 blocked · 21.3× smaller pool

Make pooling pay by moving the peaks

Correlation is a property of the schedule, not the fleet. Staggering 64 hosts leaves about 3 peaking at once, and a refused pooling case becomes one that pays.

memory.intensityLift2/2 blocked · 90×

Reach the compute already installed

Named configuration, measured gain, nothing purchased.

memory.chunkTune2/2 blocked · +60.3%

Meet the deadline the shared tier missed

Time to first token depends on the first chunk, not the whole transfer. Size to the deadline and the remainder streams behind tokens already being produced.

memory.rightSize2/2 blocked

Recover the procurement error twice

A buyable next-cycle specification at module granularity, plus the capacity additional workload can be admitted onto today.

memory.compressionPlan3/3 blocked

Cheapest sufficient compression, then stop

Applied in quality order and halted the moment it fits — never the whole stack for capacity that was not needed. Refuses when no combination reaches the target.

memory.degradedRun3/3 blocked

Finish the run on the hardware you have

When no spare device exists, “replace” means “stop”. Exclude the failed regions, finish, and replace at the maintenance window.

memory.cxlPooler+37.4%

Near and far placement, measured

250 ns to 182 ns against all-far placement.

memory.hybridPager+78.8%

Fast-tier hit rate

Against random placement across the hierarchy.

05

They measured the limit. Moving it is our job.

When a serious team publishes a hard result, they have done the field a favour. Google evaluated CXL memory pooling honestly, found the economics did not work, and published it rather than staying quiet. They were not wrong. Their measurement is correct, it is reproducible, and it is the reason anyone knows where the wall is.

A measurement tells you where the wall is. Somebody still has to get through it, and that is the job we took. Every limit this engine finds — including our own — is paired with a shipped module that moves it: the finding on the left, the lever on the right, both running, both measured at the same seed.

Where it stopsmemory.cxlEconomics

CXL pooling does not pay on a single-workload fleet — every host peaks together, so the pool must be nearly as large as the buffer it replaces.

How WUBBERY gets throughmemory.peakDecorrelate

Peaks coincide because nothing told them not to. Stagger 64 hosts across the period and about 3 peak at once instead of 64 — then the same economics call approves the pool.

21.3× smaller pool
Where it stopsmemory.bandwidthBound

Decode reaches 2.2% of installed compute. Buying more returns nothing.

How WUBBERY gets throughmemory.intensityLift

Raise arithmetic intensity instead of buying past it — we return the configuration that does it on your workload.

90× reachable compute, nothing purchased
Where it stopsmemory.rehydrateBudget

A tier that meets its target at concurrency one misses it by 38.6× at forty streams.

How WUBBERY gets throughmemory.chunkTune

Time to first token depends on the first chunk, not the whole transfer. Size the chunk to the deadline and the rest streams behind tokens already being produced.

deadline met at 40 streams
Where it stopsmemory.strandedDram

Most idle DRAM is a procurement error, not a pooling opportunity.

How WUBBERY gets throughmemory.rightSize

Then recover it twice: a buyable next-cycle specification at module granularity, and capacity additional workload can be admitted onto today.

recovered at refresh and now
Where it stopsmemory.tierPlace

The fast tiers are too small for this working set.

How WUBBERY gets throughmemory.compressionPlan

The working set is a variable. Cheapest sufficient compression, applied in quality order and stopped the moment it fits — never the whole stack for capacity you did not need.

fits inside the quality budget
Where it stopsmemory.hbmHealthDrain

Spare rows exhausted. Replace this device.

How WUBBERY gets throughmemory.degradedRun

And when no spare exists, “replace” means “stop”. Exclude the failed regions and finish the run on what remains, with the residual risk stated.

run completes, replace at the window

The pooling row is the one to read twice, and it is the clearest illustration of why the published result mattered. Peak correlation is treated everywhere as a property of the fleet — a fact about the world that decides whether an architecture is viable. It is not. It is a property of the schedule, and a schedule is something we can change. Hosts peak together because nothing ever told them not to. Nobody could have aimed at that until somebody measured precisely where the economics broke.

06

Memory fails constantly, and it stops everything

A published frontier training run recorded 72 job interruptions from uncorrectable memory errors across 54 days on 16,384 accelerators — one roughly every 18 hours. On a synchronous job, one device's memory health is a property of the entire cluster's throughput.

Devices signal before they fail: correctable errors rise, rows are remapped, and the supply of spare rows is finite. memory.hbmHealthDrain weighs the cost of draining one device against the expected rank-hours lost when the whole run stops — which is why it drains the same device before a 54-day training run and leaves it in service for a two-hour inference job. A fixed error threshold cannot tell those apart.

07

Reproduce all of it

Every figure on this page comes from one seeded command. Same seed, same answer, on your hardware, in front of you.

POST https://api.wubbery.com/v1/bench/all-modules
{ "seed": 1 }

Guards report blocked-of-attempted, never a percentage. Guards and improvements are never averaged together — a refusal and a compression ratio share no denominator. Where a figure measures exposure rather than performance, it says so.

Industry figures are published by their respective companies and researchers and are linked above. WUBBERY is independent; nothing here implies endorsement, partnership or review by any third party. Company names are used for identification only.