Hardware counters · paper Table 5
Where the cycles go.
Throughput summarizes; hardware counters explain. Read on the matcher thread over the public harness's timed loop, they price one message through Flash One and through two conforming engines from the audited field, anonymized as Projects A and B.
What one message costs
Fig. 05Cycles per message on the matcher thread, AMD EPYC Zen 5, public harness, normal scenario. The hatched share is the harness's own loop, measured with a do-nothing engine through the same entry point: 8.0 cycles batched, 12.9 through the per-message entry point Project B's adapter uses. Values are net of it.
Why the gap is that wide
Mostly work, then efficiency.
Cycles are instructions divided by the rate the core retires them. The baselines lose on both terms, and the larger share of the gap is work: for the same message they execute four to seven times as many instructions.
The two terms of the bill
Fig. 06Left, instructions per message net of the harness loop. Right, the share of the core's eight dispatch slots per cycle that retire work; hover or tap a row to see where the rest go.
27.4 cycles and 118 instructions per message, net of the harness loop, retiring about 3.3 of the eight operations the core can dispatch per cycle. A mispredict every third message, one wait on DRAM per ~1,100 messages, and code that fits the instruction cache.
It retires 15% of its dispatch slots at 1.3 instructions per cycle, misses the instruction cache about fifteen times per message, and waits on nearly three fetches from L3 and a third of a fetch from DRAM for every message.
58% of its dispatch slots go empty for want of instructions: it feeds 843 instructions per message, 816 net of the harness loop, through an instruction cache and an op cache it misses about fifteen times each.
Every counter, as published
Paper Table 5Per message, on the matcher thread. Net rows subtract the harness's own loop; the other rows include it, and report publication stays in every engine's count. Dispatch shares are of the core's eight dispatch slots per cycle; the remainder is not directly measured.
| Flash One | Project A | Project B | |
|---|---|---|---|
| Work | |||
| Cycles, total | 35.4 | 379 | 351 |
| Harness loop (null engine) | 8.0 | 8.0 | 12.9 |
| Cycles, net of the harness | 27.4 | 371 | 338 |
| Instructions, net | 118 | 473 | 816 |
| Instructions per cycle (total) | 3.71 | 1.28 | 2.40 |
| Loads + stores dispatched | 66 | 362 | 492 |
| Branches | 12.9 | 98.4 | 178 |
| Branch mispredicts | 0.31 | 1.10 | 1.69 |
| Memory | |||
| L1D demand misses | 0.28 | 4.28 | 3.55 |
| L2 misses served by L3 | 0.013 | 2.95 | 2.60 |
| DRAM demand fetches | 0.0009 | 0.32 | 0.0052 |
| dTLB page walks | 0.009 | 0.16 | 0.054 |
| Front end | |||
| I-cache misses | 0.04 | 14.9 | 15.5 |
| Op-cache misses | 1.56 | 8.44 | 15.4 |
| Dispatch slots | |||
| Retiring | 41% | 15% | 28% |
| Front-end bound | 20% | 46% | 58% |
| Back-end bound | 23% | 32% | 7% |
| Remainder (mis-speculation) | 16% | 8% | 7% |
As everywhere in the paper, these are measurements of pinned commits on one workload, not a verdict on any engine or its authors.
Counters are read through Linux's perf_event_open in user mode over exactly the public harness's timed loop, the interval its throughput is computed from: at most six events per run on the core's six programmable counters, cycles in every group as the common anchor, no event multiplexed. Across the five scenarios, Flash One's gross cost runs from 28.0 cycles (static) to 37.3 (flash-crash); on Graviton4 the same engine spends 27.1 cycles per message on the normal tape at 4.1 instructions per cycle, harness loop included.