Security
Spectre: Stealing Secrets From Speculation
Spectre is a class of hardware attacks that tricks a modern CPU into leaking secrets through the ghostly aftermath of work it never officially did. To stay fast, processors speculate — they guess which way a branch will go and race ahead executing instructions before the guess is confirmed. If the guess was wrong, the CPU quietly discards the results, so architecturally nothing happened. But that discarded work already disturbed the cache, and Spectre reads those disturbances by timing memory accesses.
What makes Spectre remarkable is that it doesn't exploit a bug in any program — it exploits a performance optimization built into virtually every high-performance chip made since the 1990s. Disclosed in January 2018 alongside Meltdown, it broke the isolation between processes, between a browser tab's JavaScript sandbox and the rest of the browser, and even between a program and its own bounds checks.
- Disclosed3 Jan 2018 (Kocher, Horn, Genkin, Yarom, et al.)
- CVEsv1 CVE-2017-5753 (bounds-check bypass), v2 CVE-2017-5715 (branch target injection)
- Leak rate~10 KB/s at >99.99% accuracy (original PoC)
- Timing gapcache hit ~tens of cycles vs DRAM ~200-300 cycles
- Probe stride4096 B (one page) per byte value → 256 distinct cache lines
- Speculation windowbounded by reorder buffer (~224 µops on Intel Skylake)
Interactive visualization
Press play, or step through manually. The visualization is yours to drive — try it before reading on.
Watch the 60-second explainer
A condensed visual walkthrough — narrated, captioned, under a minute.
Two kinds of state: what rollback forgets
The whole attack rests on a distinction the instruction set never names. A CPU keeps architectural state — the register file, memory contents, and condition flags that the ISA promises are visible to software — and, underneath it, microarchitectural state — caches, TLBs, branch-predictor tables, and buffers that exist only to make the architectural machine fast. Correctness is defined purely over the architectural layer; the microarchitecture is free to do whatever it likes as long as the visible results are the same as a simple in-order machine would produce.
To hide the hundreds of cycles a memory load or a mispredicted branch can cost, the processor executes out of order and speculatively. When it reaches a conditional branch whose outcome is not yet known, the branch predictor guesses the direction from history, and the pipeline keeps fetching and executing down the predicted path. These in-flight instructions sit in the reorder buffer (ROB), which enforces one iron rule: instructions may execute in any order, but they retire — become architecturally real — strictly in program order. If the branch resolves and the guess was wrong, every younger instruction is squashed: it never retires, the register renamer rolls back, and architecturally it is as if none of it ran.
Here is the flaw. Squashing scrubs the architectural state but leaves the microarchitectural state alone. A speculative load that pulled a line into the L1/L2/L3 cache leaves that line cached even after the instruction is discarded. The speculation was transient, but its footprint on the cache is permanent until eviction. Spectre is the art of steering that transient window over a secret and then measuring the footprint.
Spectre v1: bounds-check bypass, step by step
The canonical gadget is disarmingly ordinary:
if (x < array1_size)y = array2[array1[x] * 4096];
Read literally this is safe: the bounds check guards the array access. The attack unfolds in four moves.
1. Train the predictor. The attacker calls the gadget many times with in-bounds values of x. The conditional branch predictor learns that this branch is almost always taken (in bounds), so it will keep predicting "taken."
2. Stall the check. Now the attacker calls with a malicious, out-of-bounds x — chosen so that array1 + x points at a secret byte anywhere in the victim's address space. Crucially, the attacker first evicts array1_size from the cache, so evaluating x < array1_size must wait ~200-300 cycles for that value to arrive from DRAM.
3. Speculate over the secret. Rather than idle, the CPU trusts the trained prediction (taken) and races into the body. It speculatively computes k = array1[x] — reading the secret byte — and then issues array2[k * 4096], which pulls the cache line for index k * 4096 into cache. The multiplier of 4096 (one page) spaces the 256 possible byte values into 256 well-separated cache lines, so no two values share a line and the hardware prefetcher cannot blur them together.
4. Squash, but too late. When array1_size finally arrives, the branch resolves as not taken: the guess was wrong. The CPU squashes everything — y is never assigned, no fault is raised, the program is architecturally untouched. But the line for array2[k * 4096] is now cached, and k is the secret.
Flush+Reload: reading the cache's confession
The secret is now encoded as "which one of 256 cache lines got warmed." To read it, Spectre uses a Flush+Reload side channel over the shared array2. Beforehand, the attacker flushes all 256 probe lines from the cache with clflush. After the speculative gadget runs, the attacker times an access to each line array2[i * 4096] using a cycle counter such as rdtscp: one line will hit fast because it was speculatively cached; the other 255 miss and go to DRAM.
The gap is the entire signal. On a typical Intel core the latencies stack up as roughly L1 ~4-5 cycles, L2 ~12-14, L3 ~40-75, and main memory ~200-300+ cycles. A single threshold (say, ~150 cycles) cleanly separates a cached hit from a DRAM miss, so the fast index i reveals the secret byte value k = i. Repeat with successive out-of-bounds offsets and the attacker walks arbitrary memory, one byte at a time.
Flush+Reload needs memory shared between attacker and target (so the same physical line can be flushed and timed) — common when a library or shared page is mapped in both. Where no shared line exists, the attacker substitutes Prime+Probe: fill a target cache set with your own lines, let the victim's speculative access evict one, then detect the eviction by re-timing your lines. Either way, the cache is turned into a covert channel out of the speculative window.
The speculation window and the numbers
Everything hinges on how much work the CPU can complete before the branch resolves — the speculation window. Its width is bounded by the reorder buffer (about 224 micro-ops on Skylake-class cores) and by how long the mispredicted condition stalls. The attacker deliberately widens the window by keeping array1_size out of cache: the longer the comparison stalls, the more speculative instructions can execute, comfortably enough to read the secret and issue the probe.
How fast is it? The original 2018 proof-of-concept leaked memory at over 10 KB/s with an error rate below 0.01%, spending on the order of ~1000 cycles per byte for the speculative read plus another ~1000 for the Flush+Reload sweep. Because a single trial is noisy — a value can be prefetched, or the window can close early — practical implementations read each byte many times and take a majority vote, and add error-correcting structure so a mispredicted probe or an interrupt does not corrupt the stream. Reading a few kilobytes per second is more than enough to lift SSH keys, cookies, or password buffers out of a victim address space over seconds to minutes.
A key subtlety: Spectre v1 does not need to cross a privilege boundary at all. The gadget runs inside the victim's own code, which is precisely why it breaks language-level sandboxes: a piece of JavaScript, WebAssembly, or eBPF that is confined by bounds checks can be coerced into speculatively reading past them and leaking the host process's memory.
Spectre v2: hijacking the predictor itself
Variant 1 borrows an existing gadget in the victim. Spectre v2 (branch target injection, CVE-2017-5715) is stronger: the attacker chooses which speculative code runs. Indirect branches — a jmp *reg or a virtual-call dispatch — are predicted through the Branch Target Buffer (BTB), which on many cores is shared across contexts and indexed by branch address with limited tag bits. An attacker trains the BTB from its own address space so that a particular indirect branch is predicted to jump to an attacker-chosen target address.
When the victim later executes an indirect branch that aliases the poisoned BTB entry, the CPU speculatively jumps to a Spectre gadget the attacker selected — a short instruction sequence, already present in the victim, that reads a secret and leaks it via a cache access, exactly like the v1 body. This is spiritually close to return-oriented programming: instead of corrupting the stack, the attacker corrupts the predictor and executes gadgets purely in the transient, never-retired world. Because BTB state can leak across SMT threads and across the user/kernel boundary, v2 was the variant that most directly threatened hypervisor and kernel isolation.
Mitigations: close the window, mask the index, or fence the predictor
Because Spectre exploits a design property rather than a coding bug, there is no single patch — defenses attack different links in the chain.
- Serializing barriers (v1). Inserting an
lfenceafter a bounds check stops later loads from executing until the fence retires, collapsing the speculation window before the secret can be read. Compilers can auto-insert it (MSVC/Qspectre, GCC/Clang variants), but blanket fencing is expensive, so it is applied selectively to known gadgets. - Index masking (v1). The Linux kernel's
array_index_nospec()macro clamps the index into range with a branchless bitmask derived from the bounds check, so even mis-speculated code sees an in-bounds index. Clang's Speculative Load Hardening (SLH) generalizes this, threading a data-dependent "misspeculation" predicate into pointers so a wrong-path load reads zeros. - Retpolines (v2). Google's retpoline ("return trampoline," from Paul Turner) rewrites every indirect branch as a
call/retconstruct that steers any speculation into an inertpause; lfencespin loop instead of an attacker-poisoned target — starving branch-target injection of a landing site. - Microcode (v2). New MSR controls — IBRS/eIBRS (restrict indirect-branch speculation across privilege levels), STIBP (isolate predictors between SMT siblings), and IBPB (flush the predictor on context switch) — harden the BTB in silicon and firmware.
- Sandbox hardening. Browsers, the most exposed target, coarsened
performance.now()to ~100 microseconds and disabledSharedArrayBuffer(a high-resolution timer proxy) in 2018, then shipped Chrome Site Isolation — one renderer process per site — so cross-origin secrets simply are not in the same address space to leak.
Note that Meltdown, though disclosed together, needed a different fix: KPTI/KAISER unmaps kernel pages from user page tables, costing ~5-30% on syscall-heavy workloads. Spectre's mitigations are cheaper per site but must be reasoned about gadget by gadget.
Why it still matters
Spectre reframed a decade of processor design as a security problem. Every optimization that predicts the future — conditional and indirect branch prediction, memory disambiguation, value prediction, the return stack buffer — is a potential channel, and the family has kept growing: SpectreRSB, BranchScope, Spectre-BHB, Retbleed, and Speculative Store Bypass (v4) each abuse a different predictor. The uncomfortable truth is that architecturally correct is not the same as secure: as long as transient execution touches shared microarchitectural resources, secrets can bleed out through timing.
The practical stakes are highest for shared and sandboxed environments — multi-tenant cloud VMs, browser tabs running untrusted scripts, and kernels running unprivileged code. The long-term answer likely blends hardware that flushes or partitions predictors on domain crossings, software that keeps genuine secrets out of speculatively reachable memory, and constant-time discipline so that timing carries no data. Until then, Spectre remains a permanent reminder that speed and isolation are in tension at the very bottom of the stack.
| Attack | Microarchitectural feature abused | Boundary crossed | Primary mitigation |
|---|---|---|---|
| Spectre v1 (bounds-check bypass) | Conditional branch predictor | In-bounds vs out-of-bounds (same address space, e.g. a sandbox) | lfence barrier, array_index_nospec masking, Speculative Load Hardening |
| Spectre v2 (branch target injection) | Indirect branch predictor / BTB | Attacker process → victim's speculative control flow | Retpolines, IBRS/eIBRS, STIBP, IBPB microcode |
| Meltdown (v3) | Out-of-order execution race on permission check | User space → kernel memory (direct) | KPTI / KAISER page-table isolation |
| Common thread | Transient (squashed) instructions leave cache traces | Any isolation the CPU enforces only architecturally | Close the window and/or scrub the side channel |
Frequently asked questions
Does Spectre actually read or corrupt memory it shouldn't?
Architecturally, no. The malicious load happens only on the speculative, never-retired path, so no fault is raised and no register or memory value is changed. The secret escapes purely as a microarchitectural side effect — a warmed cache line — which the attacker then reads by timing memory accesses. Nothing in the program's visible behavior looks wrong.
How is Spectre different from Meltdown?
Meltdown (CVE-2017-5754) exploits an out-of-order race that lets user code transiently read kernel memory directly, and is largely Intel-specific; it is fixed by unmapping the kernel from user page tables (KPTI). Spectre tricks a program into leaking its own reachable memory by mistraining the branch predictor, works within a single privilege level (including sandboxes), and affects nearly all speculating CPUs — Intel, AMD, and ARM.
Why multiply the secret by 4096?
4096 bytes is one page, and it spreads the 256 possible byte values across 256 cache lines that are far enough apart that no two share a line and the hardware prefetcher won't pull in neighbors. That keeps the Flush+Reload readout unambiguous: exactly one line is warmed, and its index is the secret byte.
Can malicious JavaScript really do this in a browser?
In principle yes — that was the scariest demo. Sandboxed JS/WebAssembly runs inside the browser's address space behind bounds checks, and Spectre v1 speculatively reads past those checks. Browsers responded by coarsening timers, disabling SharedArrayBuffer, and adopting Site Isolation so each site runs in its own process, removing cross-site secrets from reach.
Can Spectre ever be fully fixed in software?
Not cleanly. Because it stems from a hardware optimization present in most CPUs, software mitigations are per-gadget (lfence, index masking, retpolines) and easy to miss, while blanket fencing is too slow to apply everywhere. Durable defense requires hardware changes plus keeping true secrets out of speculatively reachable memory.
What is the difference between Spectre v1 and v2?
v1 (bounds-check bypass) poisons the conditional branch predictor to run past a bounds check using a gadget already in the victim. v2 (branch target injection) poisons the indirect-branch/BTB predictor so the victim speculatively jumps to an attacker-chosen gadget address, giving the attacker control over which transient code runs — which is why v2's mitigations (retpolines, IBRS) target indirect branches specifically.