Optics
F-Number and Depth of Field: The Physics of Focus
Turn the aperture ring on a 50 mm lens from f/8 to f/2.8 and two things happen at once: the sensor drinks in eight times more light, and the zone of acceptable sharpness collapses from nearly two metres deep to about 60 centimetres. That single dial controls a trade-off written into geometry — the f-number N = f/D governs both how bright the image is and how much of the world falls into focus, and the two are locked together because they are the same cone of light seen from two angles.
Underneath the photographer's rules of thumb sits clean optics: similar triangles set the depth of field, an inverse-square law sets the brightness, and — at the small-aperture end — diffraction quietly caps how sharp any lens can ever be. Understanding f-number means understanding where ray optics ends and wave optics takes over.
- DefinitionN = f / D (focal length ÷ pupil diameter)
- BrightnessIlluminance ∝ 1/N²
- One stopN × √2 → half the light
- HyperfocalH = f²/(Nc)
- Diffraction spotd ≈ 2.44 λ N (Airy diameter)
- DoF scaling∝ N·c·s² / f²
Interactive visualization
Press play, or step through manually. The visualization is yours to drive — try it before reading on.
Watch the 60-second explainer
A condensed visual walkthrough — narrated, captioned, under a minute.
The f-number: one ratio, two consequences
The f-number (or f-stop) is defined as the ratio of a lens's focal length to the diameter of its entrance pupil — the aperture as seen from the front of the lens:
- N = f / D, where f is the focal length and D the effective aperture diameter.
So f/2 on a 50 mm lens means an entrance pupil of D = 50/2 = 25 mm; f/16 means D = 3.1 mm. Because N is a ratio, it is dimensionless, and — crucially — it captures light-gathering independent of focal length. A 500 mm f/4 telescope and a 25 mm f/4 microscope objective deliver the same image illuminance for a given scene brightness.
That illuminance follows an inverse-square law in N. The light collected scales with pupil area (∝ D²), while the image of a scene is spread over an area that scales with f². Combining, the illuminance on the sensor E ∝ D²/f² = 1/N². This is why the classic f-stop series — 1, 1.4, 2, 2.8, 4, 5.6, 8, 11, 16, 22 — steps by factors of √2 ≈ 1.414. Each step multiplies N by √2, so it multiplies 1/N² by exactly ½: one "stop" halves the light. The awkward numbers are just powers of √2 rounded off.
Depth of field from similar triangles
A lens focuses one object plane perfectly onto the sensor. Points nearer or farther project not to a point but to a small blur disk. The image is judged "sharp" as long as that blur is smaller than the circle of confusion c — the largest blur a viewer cannot resolve, conventionally c ≈ 0.03 mm (30 μm) for a 35 mm full-frame sensor, roughly the sensor diagonal ÷ 1500.
Geometry does the rest. A point off the focus plane sends its cone of light through the pupil of diameter D; the blur disk on the sensor has a diameter set by similar triangles between the pupil and the defocus distance. Working the geometry through the thin-lens equation gives the near and far limits of acceptable focus for a subject at distance s:
- Near limit: D_near = s(H − f) / (H + s − 2f)
- Far limit: D_far = s(H − f) / (H − s)
where H = f²/(Nc) + f is the hyperfocal distance. The depth of field is the gap D_far − D_near. Two features fall straight out of the algebra: the far limit diverges to infinity when s reaches H (the denominator H − s → 0), and for near subjects (s ≪ H) the depth of field simplifies to DoF ≈ 2Ncs²/f² — linear in N, linear in c, quadratic in subject distance, and inverse-square in focal length.
The controlling variables and their scales
The compact form DoF ≈ 2Ncs²/f² reveals every lever a photographer or engineer has:
- Aperture (N): Depth scales linearly with N. Stopping down from f/2.8 to f/16 (a factor of ~5.7 in N) deepens the zone of focus by the same factor — the reason landscapes are shot at f/11–f/16 and portraits at f/1.8.
- Subject distance (s): Quadratic. Doubling the distance quadruples the depth of field. Macro work (s of a few cm) has a depth of field of millimetres; a mountain range at s = 100 m is effectively all in focus.
- Focal length (f): Inverse-square in this near-field regime. Long lenses have shallow depth of field — but much of that is because they force you farther from the subject at a given framing; at equal image magnification the differences shrink.
- Circle of confusion (c): Set by the display and the sensor. A tiny phone sensor uses c ≈ 4–5 μm, so at the same f-number and framing it renders everything sharp — which is why phone "portrait mode" must fake background blur in software.
Worked case: a 50 mm lens at f/8, subject at 3 m, c = 0.03 mm. Then H = (0.050)²/(8 × 3×10⁻⁵) ≈ 10.5 m. Near limit ≈ 2.34 m, far limit ≈ 4.19 m, giving ≈ 1.85 m of depth of field. Open up to f/2.8 and it shrinks to about 0.60 m — exactly the trade-off in the lede.
Where geometry ends: the diffraction limit
Stopping down forever does not sharpen forever. A small aperture is a small hole, and light passing through a hole diffracts. The image of a point source is not a point but an Airy pattern — a bright central disk ringed by faint fringes. Its diameter (first minimum to first minimum) is:
- d = 2.44 λ N, for imaging at f-number N and wavelength λ.
At λ = 550 nm (green), f/2.8 gives an Airy disk of d ≈ 3.8 μm, f/8 gives ≈ 10.7 μm, and f/16 gives ≈ 21 μm. Compare those to a modern sensor's pixel pitch of ~4 μm: past about f/8–f/11 on a high-resolution full-frame sensor, the diffraction blur exceeds a pixel, and every point on the image softens. This is the diffraction limit, first quantified by George Airy in 1835 and formalized by Ernst Abbe in 1873.
So the two effects fight. Opening the aperture shrinks the depth of field but pushes back diffraction; closing it deepens focus but eventually smears fine detail. There is an optimal aperture — for most full-frame lenses around f/5.6–f/8 — where geometric aberrations have faded but diffraction has not yet dominated. This is the physics behind the well-known photographer's maxim that a lens is "sharpest in the middle of its range."
Numerical aperture, brightness, and the microscope connection
Microscopists and telescope designers use a close cousin of the f-number: the numerical aperture, NA = n·sin θ, where θ is the half-angle of the cone of light the lens accepts and n is the refractive index of the medium (1.0 in air, ~1.52 in immersion oil). For a lens focused far away the two are related by N ≈ 1/(2·NA). Higher NA means a fatter cone, so both resolution and light-gathering improve — the Rayleigh resolution limit is δ ≈ 0.61 λ / NA.
This unifies the fields. Whether you call it N or NA, the same wide cone of light that gathers photons quickly is also the cone that, in ray optics, produces a big defocus blur — hence shallow depth of field — and, in wave optics, produces a small Airy disk — hence high resolution. Light-gathering, depth of field, and resolution are three faces of the same geometry of the marginal ray. You cannot independently maximize all three; the f-number is the knob that trades among them.
Consequences show up everywhere: a fast f/1.2 portrait lens isolates a face with creamy background blur but is unforgiving of focus error (mm-thin depth of field on the eye); an f/22 macro shot of an insect keeps the whole body sharp but loses fine texture to diffraction and needs a bright flash to fight the 1/N² light loss.
T-stops, bokeh, and other subtleties
Several refinements separate the textbook f-number from what a real lens does:
- T-stops. The geometric f-number ignores that lens elements absorb and reflect a few percent of light. Cinematographers use the T-stop — T = N/√(transmission) — which measures actual transmitted light, so exposures match across different lenses. A lens marked f/2.0 might be T2.2.
- Bokeh. The out-of-focus blur is not just a size but a shape — the image of the aperture itself. A polygonal diaphragm renders point highlights as pentagons or heptagons; rounded blades give circular "bokeh balls." This is literally the pupil function projected onto the sensor.
- Bellows factor. N = f/D assumes focus at infinity. At close focus the image distance grows, dimming the image; the effective f-number is N_eff = N(1 + m), where m is magnification. At 1:1 macro (m = 1) you lose two full stops of light.
- The c convention is not physics. Depth of field is a perceptual criterion, not a hard boundary. Change the print size, viewing distance, or how picky the viewer is, and c changes — so the "in focus" zone is genuinely fuzzy at its edges. Only one plane is ever truly in focus.
| Quantity | f/2.8 (wide) | f/16 (narrow) | Why |
|---|---|---|---|
| Pupil diameter D | 17.9 mm | 3.1 mm | D = f/N |
| Relative light gathered | 1× (reference) | 1/33 (≈5 stops less) | ∝ 1/N² |
| Hyperfocal distance H | 29.8 m | 5.3 m | H = f²/(Nc) |
| Depth of field at 3 m | ≈ 0.60 m | ≈ 5.0 m (1.9 m → 6.9 m) | ∝ N near, diverges only past H |
| Airy disk diameter | 3.8 μm | 21 μm | d ≈ 2.44 λN — diffraction grows |
Frequently asked questions
Why does a smaller aperture give more depth of field?
A smaller aperture (larger N) narrows the cone of light converging to each image point. Because the cone is thinner, an out-of-focus point spreads into a smaller blur disk on the sensor, so it stays under the circle-of-confusion threshold over a longer range of distances. Geometrically the blur diameter scales with the pupil diameter D = f/N, so halving D via a two-stop change roughly halves the near-field blur and doubles the depth.
Why does each f-stop change light by a factor of two, not a factor of √2?
The f-stop numbers step by √2 in N, but image brightness goes as 1/N². Squaring √2 gives exactly 2, so each labeled stop halves (or doubles) the light. The √2 in the aperture series is the diameter change needed to change the pupil area by a factor of two, since area scales with diameter squared.
How small an aperture is too small?
When the diffraction-limited Airy disk d ≈ 2.44λN grows larger than your sensor's pixel pitch, extra depth of field comes at the cost of overall sharpness. At green light (550 nm), f/8 gives a ~10.7 μm disk and f/16 gives ~21 μm — both exceed a ~4 μm pixel, so a high-resolution full-frame camera visibly softens past roughly f/8–f/11. This is the diffraction limit, and it caps every lens no matter how well made.
Why do phone cameras keep everything in focus while big cameras blur backgrounds?
Depth of field scales as ≈ 2Ncs²/f², and phones use tiny sensors with very short focal lengths (f ≈ 4–6 mm) and small circles of confusion (~4 μm). The short f gives an enormous depth of field even wide open, so almost the whole scene is sharp. That is why phones must simulate background blur computationally in 'portrait mode' rather than producing it optically.
What is the hyperfocal distance and why is it useful?
The hyperfocal distance is H = f²/(Nc) + f, the focus distance at which the far limit of depth of field just reaches infinity. Focus there and everything from H/2 out to infinity is acceptably sharp — the deepest possible zone of focus for that aperture. Landscape photographers use it to get both foreground and horizon sharp in one shot; for a 50 mm lens at f/16, H ≈ 5.3 m, so focusing at 5.3 m keeps 2.6 m to infinity sharp.
Is depth of field a sharp boundary?
No. Only a single object plane is ever perfectly focused; sharpness degrades continuously on either side. The 'depth of field' is simply the range where the blur disk stays smaller than a chosen circle of confusion, which itself depends on print size, viewing distance, and viewer acuity. Change any of those and the boundaries move — the edges of focus are inherently a soft perceptual convention, not a physical wall.