Technical Note
Benchmarking memory-bound inference is easy to get wrong. Peak bandwidth figures rarely predict tokens per second per user, because the number that matters is sustained bandwidth under a realistic mix of request lengths and batch sizes.
Methodology
We hold model, precision, and batch policy fixed, then vary only the memory attachment. Each run reports sustained bandwidth, tail latency at the 99th percentile, and tokens per second per user, averaged across a fixed replay of production traffic shapes.
Reproducing the numbers
The harness, traffic shapes, and raw run logs are available to design partners. We publish the full configuration alongside every figure so results can be checked against an equivalent HBM-attached system.
Want to go deeper?
Talk to our team about optical memory for your fleet.
Get in touch