Did you ever wonder how an LLM generates one token?
A distilled understanding of what happens when an LLM generates one token, with an interactive explainer.
Read the field note →MEASURED / FIELD NOTES FROM THE OPERATOR’S SIDE
I test how production AI performs, scales, fails, and costs from the user/operator's side.
I'm an enterprise platform engineer with years of experience; these are my own views.
I publish one measured report a month on AI systems, their infrastructure, and the decisions they shape.
01 / THE THESIS
Operators need a different instrument: freeze the hardware, sweep the workload, and measure the system at the point where it will actually run.
02 / LATEST FIELD NOTE
A distilled understanding of what happens when an LLM generates one token, with an interactive explainer.
Read the field note →03 / IN THE LAB
INSTRUMENT 01 / LIVE
Move the Cassandra mix from reads to writes. Compare a simple capacity model with three measured runs—and see exactly where the claim stops.
Open the Mix-Axis Explorer →MODELLED CAPACITY ↓
No noise. Unsubscribe at any time.