Instruction latency analysis usually focuses on performance
optimization—making code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of
single-instruction performance.
Strategy: Use fxrstor64 to load 512-byte FPU/MMX/XMM state from a
high-latency MMIO region in the PCIe fabric, then starve the fabric while the
load is in flight — a fleet of hammer cores pounds a different high-latency
MMIO register with tight 4-byte reads, saturating the PCIe root complex and
endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must
queue behind all that contending traffic.
Contender: AMD Ryzen 7 5800H
; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi
; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax
🏆 Score: 198,002,498,236 cycles
🏆 Time: 62 seconds
A spec-violating unaligned ymm0 load that forced non-posted dword transactions from stalled GPU registers was used to break the fundamental design of System Management Mode in smiiiiiiiiiiiiiiii.
vmovdqu 0xfcc003b1, %ymm0- Instructions may use whatever setup is necessary, but only a single instruction is eligible to be scored.
- Trapped/emulated/virtualized instructions may only time the trap, not the handler.
- Instructions must not be interruptible. rep movs,pause, etc. are disqualified.
- Times are normalized based on the CPU base clock frequency.
- All platforms must be in their factory stock configurations - no hardware modifications.
Strategy: nop does nothing. It opens the leaderboard accordingly.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
nop
Score: 1 cycles
Time: 0 nanoseconds
Strategy: Regular nop was too short, but how do we make nothing take
longer? Try a lonnnnnng nop.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
data16 data16 data16 data16 data16 data16 data16 nopl 0x00000000(%%eax,%%eax,1)
Score: 20 cycles
Time: 7 nanoseconds
Strategy: Just a reference instruction to get our bearings.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
rdtscScore: 49 cycles
Time: 18 nanoseconds
Strategy: Use 128-bit dividend (rdx:rax=2:0) with small divisor to push the
quotient above the ceiling imposed by sign-extension, driving the longest path
through the divider microcode.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
xorq %rax, %rax ; rax = 0 (low 64 bits of dividend)
movq $2, %rdx ; rdx = 2 (high 64 bits: full dividend = 2^65)
movq $5, %rbx ; divisor → quotient = 2^65/5 ≈ 7.4×10^18
idivq %rbx
Score: 77 cycles
Time: 28 nanoseconds
Strategy: Use maximum nesting depth (31) to force 30 display-pointer loads
and pushes through the microcode display-walk path.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
enter $0, $31 ; 0 bytes allocated, nesting depth 31 (maximum)Score: 112 cycles
Time: 41 nanoseconds
Strategy: Try a small denormal to trigger an FP microcode assist.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x0000000000000001, %rax
movq %rax, -8(%rsp)
fldl -8(%rsp)
Score: 133 cycles
Time: 49 nanoseconds
Strategy: Just ensure the cache line is dirty.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
clflush (%rax) ; rax -> dirty cache line resident in L3Score: 165 cycles
Time: 60 nanoseconds
Strategy: Use exponent 0x7ff to reach 'special value' processing in
microcode; positive/negative, NaN/inf doesn't seem to make a difference, go with
QNaN.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x7fffffffffffffff, %rax
movq %rax, -8(%rsp)
fldl -8(%rsp)
fsin
Score: 257 cycles
Time: 94 nanoseconds
Strategy: Saturate all write-combining line-fill buffers with movnti
stores to distinct cache lines, forcing mfence to drain the full LFB write
path to the uncore before retiring.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movnti %r9, 0*64(%rdi) ; ×16 distinct cache lines — saturate the write-combining LFBs
; …
movnti %r9, 15*64(%rdi)
mfence ; must drain all pending LFB writes before retiring
Score: 326 cycles
Time: 120 nanoseconds
Strategy: Nothing for now, just check how long it takes to invalidate the
TLB.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov %rax, %cr3Score: 352 cycles
Time: 110 nanoseconds
Strategy: Hit x87 FP microcode assist path by using denormal source operand.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
fldl subnorm ; 1e-310: value < DBL_MIN, biased exponent = 0
faddl subnorm ; source is subnormal → FP microcode assist
Score: 677 cycles
Time: 249 nanoseconds
Strategy: Align lock-prefixed operand to straddle cache-line
boundary, forcing CPU to assert the external bus lock rather than using the fast
MESI cache-coherence path.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
; split_ptr % 64 == 63 — dword spans bytes 63 (line N) and 64–66 (line N+1)
lock xaddl %r9d, (%rdi)
Score: 865 cycles
Time: 319 nanoseconds
Strategy: Use subnormal divisor, hardware hands control to microcode
assist, assist normalizes operand, performs the division, then restores
architectural state.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x3ff0000000000000, %rax ; 1.0 (normal dividend)
movq %rax, -8(%rsp)
fldl -8(%rsp) ; ST(0) = 1.0
movabsq $0x0000002000000000, %rax ; 6.79e-313 (subnormal divisor)
movq %rax, -8(%rsp)
fdivl -8(%rsp) ; ST(0) = 1.0 / subnormal → FP assist
Score: 883 cycles
Time: 325 nanoseconds
Strategy: Use rakefield to find
the highest latency CPUID leaves.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movl $6, %eax
cpuid
Score: 1248 cycles
Time: 460 nanoseconds
Strategy: Execute in a tight loop to deplete the hardware entropy pool faster than it can be refilled, forcing subsequent calls to stall while the entropy source recovers.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
rdrand %raxScore: 5,579 cycles
Time: 2.057 microseconds
Strategy: Use project:nightshyft
to identify high latency MSRs. MCG_CTL on Zen look like a winner: may be a
microcode quiesce and synchronize on MCA error banks across hardware units, some
potentially off-die, requiring fabric-level communication rather than a simple
local register write.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
movl $0x17b, %ecx ; MCG_CTL
wrmsr
Score: 34,304 cycles
Time: 10.742 microseconds
Strategy: Target an I/O port that straddles a NIC device register boundary,
triggering the device to quiesce its TX DMA engine on each write.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov $0xf019, %dx
outl %eax, %dx
Score: 49,857 cycles
Time: 15.580 microseconds
Strategy: Use project:nightshyft
to identify high latency model-specific-registers: VIA uses an undocumented
register at 0x133 that gives wildly high response time. No idea what it does.
Contender: VIA Eden Processor 800MHz
movl $0x133, %ecx ; undocumented MSR
rdmsr
Score: 161,602 cycles
Time: 202.004 microseconds
Strategy: Fully load L1/L2/L3 caches with dirty lines to force DRAM
writeback of entire hierarchy.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
wbinvdScore: 1,616,480 cycles
Time: 506.165 microseconds
Strategy: Target I/O port mapped to an ACPI PM block where an unaligned
4-byte read decodes into multiple non-posted loads from wherever this port goes.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov $0x0413, %dx
inl %dx, %eax
Score: 12,524,415 cycles
Time: 3.921769 milliseconds
Strategy: Use mmiotic to identify
high-latency deadspace in PCIe fabric, hit unkown GPU register.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
movl 0xfcc003b0, %esiScore: 443,937,696 cycles
Time: 139.010268 milliseconds
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit
unknown GPU register, use 8-byte MMIO read to get two dword register accesses,
which isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
movq 0xfcc003b0, %raxScore: 887,716,864 cycles
Time: 277.971228 milliseconds
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit
unknown GPU register, use 16-byte MMIO read to get four dword register accesses,
which isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b0, %xmm0Score: 1,774,555,776 cycles
Time: 555.664133 milliseconds
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit
unknown GPU register, use 32-byte MMIO read to get eight dword register
accesses, which still isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b0, %ymm0Score: 3,549,079,296 cycles
Time: 1.111345034 s
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit
unknown GPU register, use 32-byte unaligned MMIO read to get nine dword
register accesses, which is even less allowed than the aligned version, but
works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b1, %ymm0Score: 4,453,212,256 cycles
Time: 1.394428818 seconds
Strategy: Use mmiotic to identify
high-latency deadspace in PCIe fabric, isolate region near 0's and offset state
to avoid MXCSR corruption (possibly VGA buffer?), use fxrstor64 to load 512-byte
FPU/MMX/XMM state from MMIO, forcing CPU to process 512 bytes of I/O
transactions through slowest available memory aperture.
Contender: AMD Ryzen 7 5800H
movl $0xfcc68830, %rsi
fxrstor64 %rsi
Score: 74,584,168,512 cycles
Time: 23.354502677 seconds
Strategy: Extend fxrstor64 (baseline) by starving the fabric while the
load is in flight — a fleet of hammer cores pounds a different high-latency
MMIO register with tight 4-byte reads, saturating the PCIe root complex and
endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue
behind all that contending traffic.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
; CPU 0 — timed instruction
movl $0xfcc68830, %rsi
fxrstor64 %rsi
; CPUs 1..N — hammer loop against a different high-latency location
movl 0xfcc68858, %eax
🏆 Score: 198,002,498,236 cycles
🏆 Time: 62 seconds
Strategy: Leverage extended AVX state in Sapphire Rapids with MMIO approach
from fxrstor64: xsave state area is 8KB vs 512 bytes, 16x size ->
1,000,000,000,000 cycles
Contender: TODO
; XCR0 must enable AMX components (bits 17-18); state area ~8KB
xrstor64 (%rsi) ; rsi -> MMIO region, same technique as fxrstor64
- T.B.D.
- T.B.D.
The assembly hall-of-shame is a research effort from Christopher Domas (@xoreaxeaxeax).