Security

Control-Flow Integrity: Stopping Hijacked Jumps

Control-Flow Integrity (CFI) is a defense that forces a running program to follow only the jumps, calls, and returns its own source code allows — so an attacker who corrupts a code pointer cannot bend execution into an arbitrary place. Modern exploits rarely inject new code; memory is marked non-executable, so instead they overwrite a return address or a function pointer and stitch together snippets of the program's own instructions ("gadgets") into a malicious sequence. CFI kills that trick by checking, at every indirect branch, that the target belongs to a precomputed set of legitimate destinations drawn from the program's control-flow graph. A pointer aimed at the middle of a gadget fails the check and the program aborts instead of being hijacked.

  • IntroducedAbadi, Budiu, Erlingsson & Ligatti, CCS 2005
  • ProtectsIndirect calls/jumps (forward) + returns (backward)
  • Check costO(1) per indirect transfer
  • OverheadLLVM CFI ~1%; original Abadi CFI ~16% avg (45% max)
  • HardwareIntel CET (IBT+shadow stack, 2020), ARM BTI/PAC
  • Backward edgeShadow stack — trusted 2nd copy of return addresses

Interactive visualization

Press play, or step through manually. The visualization is yours to drive — try it before reading on.

Open visualization fullscreen ↗

Watch the 60-second explainer

A condensed visual walkthrough — narrated, captioned, under a minute.

The threat: code reuse when you cannot inject code

Two decades ago a stack overflow let an attacker write shellcode onto the stack and jump to it. The universal countermeasure — W^X / DEP (Data Execution Prevention: pages are either writable or executable, never both) — made injected bytes non-runnable. Attackers responded not by defeating W^X but by reusing the program's existing executable code. The building block is a gadget: a short instruction sequence already present in the binary or a shared library that ends in an indirect transfer.

In Return-Oriented Programming (ROP), each gadget ends in ret. The attacker uses a memory-corruption bug to plant a chain of return addresses on the stack; every ret pops the next address and executes the next gadget, so a corrupted stack becomes a little program. Hovav Shacham's 2007 paper showed that the gadgets in a standard C library form a Turing-complete instruction set — an attacker who controls the stack can compute anything. Jump-Oriented Programming (JOP) does the same using gadgets ending in indirect jmp/call and a dispatcher gadget, sidestepping any defense that only watches returns.

The common thread: the attacker never writes new code. They overwrite a code pointer — a saved return address, a C function pointer, a C++ vtable pointer, or a GOT entry — so that a legitimate indirect branch lands somewhere the programmer never intended. CFI attacks that root cause directly.

The core idea: constrain every indirect transfer to the CFG

Instructions that change the program counter come in two flavors. Direct branches encode their destination as an immediate constant inside the instruction; because code pages are read-only, that target is immutable and can never be hijacked. Indirect branches take their destination from a register or memory location — an indirect call rax, an indirect jmp [rbx], or a ret that reads its target off the stack. These are exactly the transfers an attacker can redirect, so these are exactly the ones CFI instruments.

Ahead of time, the compiler or a binary analyzer computes the control-flow graph (CFG): the static set of edges the program is allowed to take. CFI then enforces one invariant at runtime:

  • At every indirect transfer, the actual target must lie in the set of CFG-permitted successors for that site.

The original 2005 formulation by Abadi, Budiu, Erlingsson, and Ligatti implements this with labels: assign an identifier to each valid destination, embed that identifier as a marker at the destination, and insert a check before each indirect branch that reads the marker at the computed target and compares it to the expected label. A match means the target is a sanctioned entry point; a mismatch — a pointer aimed into the middle of a gadget, where no valid label sits — triggers an immediate abort. Each check is a handful of instructions and runs in O(1); the cost is a small constant added to every indirect branch, not to straight-line code.

Crucially, computing the exact target set of an indirect call is equivalent to precise points-to analysis, which is undecidable in general. Every practical CFG is therefore a conservative over-approximation: it may permit a few edges that never actually occur. How tightly it approximates — the precision — is what separates strong CFI from weak CFI.

Forward edge: type checks, bitmaps, and hardware landing pads

The forward edge covers indirect calls and jumps. Three enforcement styles dominate real deployments, trading precision against cost:

  • Type-based (Clang/LLVM CFI). -fsanitize=cfi-icall groups functions by their type signature and emits a compact jump table per type. An indirect call through a void(*)(int) pointer is allowed only to reach a function of that exact type, checked with a fast range-and-alignment test against the sorted table. This is type-precise forward CFI, but it needs whole-program visibility — Link-Time Optimization (LTO), or the cross-DSO extension for shared libraries. Measured overhead on SPEC is around 1%. Chrome and the Android platform ship it.
  • Bitmap (Microsoft Control Flow Guard). Windows CFG maintains a per-process bitmap with roughly one bit per 16-byte-aligned code granule; a set bit marks a valid indirect-call target. The compiler inserts a call to _guard_check_icall before every indirect call, which indexes the bitmap and faults if the bit is clear. CFG is coarse: it permits any valid function entry, not the specific ones for that call site, so a same-signature-agnostic attacker still has many targets. Its successor XFG adds a type hash to tighten the set.
  • Hardware landing pads (Intel CET-IBT, ARM BTI). Intel's Indirect Branch Tracking requires that the instruction landing after any indirect call/jump be an ENDBR64; the CPU enters a wait-for-endbranch state and raises a control-protection fault (#CP, vector 21) otherwise. ARMv8.5 BTI works the same way with a BTI instruction. The marker is a 4-byte no-op on older chips, so binaries stay compatible, and enforcement is essentially free — but it is coarse, since any landing pad is a legal target.

Backward edge: shadow stacks and pointer authentication

The backward edge — the ret instruction — is both the most attacked (it is the heart of ROP) and the one where CFI can be made exact. A static CFG can only say "function f may return to any of its N callers," which still leaves an attacker a menu of call-preceded gadgets. The fix is to record where control actually came from, giving context sensitivity no static set can match.

A shadow stack is a second, protected stack holding a trusted duplicate of every return address. On each call, the return address is pushed to both the ordinary stack and the shadow stack; on each ret, the two top values are compared. If a memory-corruption bug altered the return address on the normal stack, it no longer matches the shadow copy, and the program aborts. Intel CET implements this in hardware (Tiger Lake and AMD Zen 3, 2020): the CPU keeps a shadow-stack pointer (SSP), pages are tagged "shadow stack" so ordinary stores cannot touch them, and ret faults on any mismatch. Software shadow stacks cost a few percent; the hardware version is near-zero.

ARM Pointer Authentication (PAC, ARMv8.3) takes a cryptographic route instead of a second stack. It computes a keyed MAC — a Pointer Authentication Code — over the pointer value plus a 128-bit secret key and a context modifier (the stack pointer, for return addresses), and stuffs the MAC into the unused high bits of the 64-bit pointer. PACIASP signs the return address on entry; AUTIASP verifies it before ret. Corrupt the pointer and the MAC no longer validates, so using it faults. Apple's arm64e (A12, 2018) deploys PAC widely. The catch: the PAC field is small (often ~16 bits), so it is theoretically brute-forceable, and speculative attacks such as PACMAN (2022) have probed it.

Coarse vs fine-grained, and how the guarantees break

CFI strength is measured by how small the permitted target sets are — how few equivalence classes the indirect transfers collapse into. Coarse-grained CFI uses a handful of classes: "any function entry" for calls, "any instruction after a call" for returns. That is cheap and easy to deploy but leaves large gadget pools. Research systems Out of Control (2014) and Losing Control / Stitching the Gadgets (2014–15) built working ROP chains that stayed entirely inside such coarse policies. Fine-grained CFI shrinks the classes toward the true CFG — per-type or per-callsite — and is dramatically harder to bypass.

But even ideal static CFI has ceilings. Control-Flow Bending (Carlini et al., USENIX Security 2015) showed that with a corruptible backward edge, an attacker can "bend" execution along edges that are individually legal in the CFG yet compose into an exploit — proving that a precise shadow stack for returns is essential, not optional. Counterfeit Object-Oriented Programming (COOP) chains whole existing C++ virtual functions by forging fake objects, so every call lands on a legitimate vtable entry and defeats forward-edge CFI that isn't vtable-precise. These results also cautioned against the Average Indirect-target Reduction (AIR) metric: a policy can report 99% AIR and still be exploitable, because the residual 1% of targets is enough. The field moved to counting actual usable gadgets and to demanding exact backward-edge protection.

Two assumptions must also hold or the whole edifice falls: code must be non-writable (W^X), otherwise an attacker rewrites the checks themselves; and the CFI metadata — label tables, the CFG bitmap, the shadow stack — must be beyond the attacker's write primitive, which is why hardware tags the shadow stack and PAC hides keys in privileged registers.

Costs, deployment, and how it is measured

The reason CFI is now ubiquitous is that its cost fell to almost nothing. The 2005 prototype added around 16% average overhead (up to 45%) on SPEC CPU; software forward-edge CFI in LLVM is about 1%; and hardware IBT and shadow stacks are effectively free because the checks fold into the CPU pipeline. Overhead is benchmarked the classic way — run SPEC CPU with and without instrumentation — while security is measured by gadget-reduction counts and by whether known exploit techniques (ROP chains, COOP, control-flow bending) can still be constructed against the enforced policy.

Real deployments now blanket the stack. Clang/LLVM CFI hardens Chrome and Android's userspace and kernel. Microsoft Control Flow Guard (and XFG) ships across Windows, with Intel CET surfaced as "Hardware-enforced Stack Protection." Intel CET (IBT + shadow stack) and AMD's shadow stack are in shipping silicon and supported by the Linux kernel and glibc. ARM BTI and PAC protect Apple platforms, Android, and Linux on ARMv8.3+. grsecurity's RAP offers cryptographic return-address and type-based forward CFI for hardened Linux kernels.

CFI does not stand alone — it is one layer alongside ASLR (which randomizes where gadgets live), stack canaries (which catch contiguous stack overwrites), and sandboxing. What CFI adds is a hard, deterministic invariant: no matter which pointer an attacker corrupts, control can only ever flow where the program's own graph already permitted, and a hijacked jump becomes a crash instead of a compromise.

Deployed CFI mechanisms: which edge they guard, how, and at what cost
MechanismEdge protectedHow it enforcesGranularity / cost
Abadi et al. CFI (2005)BothInsert label at targets; compare label before indirect branchFine-ish; ~16% avg overhead
Clang/LLVM CFI (-fsanitize=cfi)ForwardPer-type jump tables; range/bitset check on callType-precise; ~1% overhead (LTO)
Windows Control Flow GuardForwardGlobal bitmap of valid call-target addressesCoarse (any valid entry); low overhead
Intel CET — IBT + Shadow StackBothENDBR64 landing pads; CPU-maintained shadow stack, RET comparesCoarse forward, exact backward; ~0% overhead
ARM BTI + Pointer AuthenticationBothBTI landing pads; cryptographic MAC (PAC) in pointer bitsCoarse forward, crypto backward; near-zero

Frequently asked questions

Why isn't W^X / DEP enough on its own?

DEP stops an attacker from executing bytes they wrote into data pages, which defeats classic shellcode injection. But it does nothing about reusing the program's own executable code. Return-oriented and jump-oriented programming chain existing instruction fragments together by corrupting code pointers, so execution stays inside legitimate executable pages the whole time. CFI is the layer that constrains where those pointers may point.

What is the difference between forward-edge and backward-edge CFI?

Forward-edge CFI protects indirect calls and jumps — it checks that a corrupted function pointer or vtable pointer still lands on a valid function entry, using type checks, bitmaps, or landing-pad instructions like ENDBR/BTI. Backward-edge CFI protects returns, which are the core of ROP, usually with a shadow stack that keeps a trusted second copy of every return address. Strong CFI needs both, because attacks target both edges.

Why is a shadow stack more precise than a static CFG for returns?

A static graph can only say a function may return to any of its callers, so it still permits a set of targets an attacker can pick from. A shadow stack records the exact address the current call came from and demands the return match it, giving per-invocation context that no static set can express. That is why Control-Flow Bending showed a precise shadow stack is essential rather than optional.

What makes coarse-grained CFI weaker than fine-grained CFI?

Coarse-grained CFI lumps many destinations into a few equivalence classes — for example, allowing an indirect call to reach any function entry rather than only same-type functions for that specific call site. Larger permitted sets leave more usable gadgets, and researchers have built working ROP chains that never leave a coarse policy. Fine-grained CFI shrinks the sets toward the true control-flow graph, greatly reducing the attacker's options.

What does the ENDBR64 instruction do in Intel CET?

ENDBR64 is a landing-pad marker for Indirect Branch Tracking. After any indirect call or jump, the CPU enters a state where the very next instruction must be an ENDBR64; if it is anything else, the processor raises a control-protection fault and the program aborts. It is a coarse forward-edge check — any ENDBR is a legal target — but the marker is a harmless no-op on older CPUs, so it costs nothing and preserves compatibility.

How does ARM Pointer Authentication protect a return address without a shadow stack?

PAC computes a keyed cryptographic MAC over the pointer plus a secret key and a context value (the stack pointer for return addresses) and stores that MAC in the pointer's unused high bits. Signing happens on function entry and verification before the return; if memory corruption changed the pointer, the MAC no longer validates and dereferencing it faults. The tradeoff is a small MAC field, which is why brute-force and speculative attacks like PACMAN are a concern.