Flash One

Hardware counters · paper Table 5

Where the cycles go.

Throughput summarizes; hardware counters explain. Read on the matcher thread over the public harness's timed loop, they price one message through Flash One and through two conforming engines from the audited field, anonymized as Projects A and B.

What one message costs

Fig. 05

Cycles per message on the matcher thread, AMD EPYC Zen 5, public harness, normal scenario. The hatched share is the harness's own loop, measured with a do-nothing engine through the same entry point: 8.0 cycles batched, 12.9 through the per-message entry point Project B's adapter uses. Values are net of it.

Cycles per message, net of the harnessfewer is better
Flash One27.4≈ 6.1 ns at 4.51 GHz
Project A37113.5× the cycles
Project B33812.3× the cycles
0100200300400
Harness loop, measured with a do-nothing engineThe engine's own work

Why the gap is that wide

Mostly work, then efficiency.

Cycles are instructions divided by the rate the core retires them. The baselines lose on both terms, and the larger share of the gap is work: for the same message they execute four to seven times as many instructions.

The two terms of the bill

Fig. 06

Left, instructions per message net of the harness loop. Right, the share of the core's eight dispatch slots per cycle that retire work; hover or tap a row to see where the rest go.

Instructions per messagefewer is better
Flash One118
Project A4734.0×
Project B8166.9×
0300600900
Dispatch slots retiringmore is better
Flash One41%
Project A15%
Project B28%
0%25%50%75%100%
Flash One

27.4 cycles and 118 instructions per message, net of the harness loop, retiring about 3.3 of the eight operations the core can dispatch per cycle. A mispredict every third message, one wait on DRAM per ~1,100 messages, and code that fits the instruction cache.

Project A · both sides of the core

It retires 15% of its dispatch slots at 1.3 instructions per cycle, misses the instruction cache about fifteen times per message, and waits on nearly three fetches from L3 and a third of a fetch from DRAM for every message.

Project B · the front end

58% of its dispatch slots go empty for want of instructions: it feeds 843 instructions per message, 816 net of the harness loop, through an instruction cache and an op cache it misses about fifteen times each.

Every counter, as published

Paper Table 5

Per message, on the matcher thread. Net rows subtract the harness's own loop; the other rows include it, and report publication stays in every engine's count. Dispatch shares are of the core's eight dispatch slots per cycle; the remainder is not directly measured.

Flash OneProject AProject B
Work
Cycles, total35.4379351
Harness loop (null engine)8.08.012.9
Cycles, net of the harness27.4371338
Instructions, net118473816
Instructions per cycle (total)3.711.282.40
Loads + stores dispatched66362492
Branches12.998.4178
Branch mispredicts0.311.101.69
Memory
L1D demand misses0.284.283.55
L2 misses served by L30.0132.952.60
DRAM demand fetches0.00090.320.0052
dTLB page walks0.0090.160.054
Front end
I-cache misses0.0414.915.5
Op-cache misses1.568.4415.4
Dispatch slots
Retiring41%15%28%
Front-end bound20%46%58%
Back-end bound23%32%7%
Remainder (mis-speculation)16%8%7%

As everywhere in the paper, these are measurements of pinned commits on one workload, not a verdict on any engine or its authors.

Counters are read through Linux's perf_event_open in user mode over exactly the public harness's timed loop, the interval its throughput is computed from: at most six events per run on the core's six programmable counters, cycles in every group as the common anchor, no event multiplexed. Across the five scenarios, Flash One's gross cost runs from 28.0 cycles (static) to 37.3 (flash-crash); on Graviton4 the same engine spends 27.1 cycles per message on the normal tape at 4.1 instructions per cycle, harness loop included.