Captions for short-form: styling, timing and the errors that cost you
Most short-form is watched muted, so captions are the primary text layer, not an accessibility afterthought. How to style and time them so they help retention.
Short answer
Captions should sit in the middle third of the frame, run one to three words per beat rather than full sentences, use a heavy sans-serif with a hard outline or solid background, and emphasise the two or three words that carry the point. Auto-generated captions need human correction, because a single wrong name or number visibly breaks trust.
Captions on short-form are not an accessibility feature that happens to be visible. On a talking-head clip they are the primary text layer, read by most of the audience, and their styling changes retention.
Here is what holds up.
Placement
Middle third. Roughly centre, or a little below. Never the bottom.
Every major platform covers the bottom of the frame with its own interface: the post caption, the username, the sound attribution, the action buttons. Captions placed in the natural subtitle position get buried under all of it, and you will not see this in your editor.
Check your placement against a real screenshot of each platform, not against the raw video.
Chunking and timing
One to three words per beat. Not single words, which feel frantic across a forty-second clip. Not full sentences, which force reading rather than glancing.
Change on the speech rhythm. Captions should turn over with the natural beats of the sentence. Fixed-interval switching feels mechanical because it fights the delivery.
Slightly early is better than slightly late. A caption arriving a frame or two before the word feels natural. Arriving after feels like lag, and the viewer notices immediately even if they cannot say why.
Hold the last chunk. Do not clear the final words the instant the audio ends. A short hold lets the point land.
Styling
Heavy weight sans-serif. Legibility at small sizes on a moving background. Condensed faces fit more per line and stay readable.
Hard contrast. A thick outline, a drop shadow, or a solid block behind the text. Video is a moving background and text without a contrast treatment will disappear over something light.
Large. Larger than looks right on a desktop editor. Preview on an actual phone before deciding.
Emphasis on two or three words (the same principle as a hook). Colour or scale on the words carrying the point. This is the single highest-value caption decision, because it tells a skimming viewer what the clip is about without them hearing a word.
Restraint on animation. Some movement helps. Every word bouncing and rotating pulls attention away from the meaning and now reads as dated.
Accuracy
Auto-transcription is good and not good enough to ship unchecked.
It reliably fails on:
- Proper nouns: people, companies, products, places
- Numbers, particularly currency and large figures
- Technical vocabulary and jargon
- Accents, and any speaker who talks quickly
- Homophones that change the meaning entirely
A single visible error, a misspelled guest name or a wrong figure, is noticed instantly and it undermines everything else in the clip. The correction pass is short. Skipping it is the most common visible quality gap between clips that look professional and clips that do not.
The things that go wrong
Captions under the interface. Covered above and still the most common error.
Too small. Edited on a large monitor, watched on a phone.
Low contrast. White text over a light wall, invisible for half the clip.
Too much on screen. Four lines of text is a paragraph, not a caption.
Punctuation clutter. Full stops and commas add visual noise without adding readability at this chunk length. Question marks earn their place. Most other punctuation does not.
Inconsistency across clips. Caption styling is a brand asset. Someone should be able to recognise your clips from the captions alone, which only happens if the style is fixed and reused.
A default that works
If you need a starting point rather than a system:
- Position: centred, at roughly 45 percent from the top
- Chunking: two to three words, switching on speech beats
- Font: heavy or extra-bold condensed sans-serif
- Size: large enough that three words fill most of the frame width
- Treatment: white text, thick dark outline
- Emphasis: one accent colour on two or three words per clip
- Animation: a simple cut or a fast fade, nothing per-word
- Accuracy: every clip read through by a person before export
Fix that once, apply it to everything, and revisit it in six months rather than every clip.
Frequently asked questions
Do platform auto-captions count?
They help accessibility and they do not do the job burned-in captions do. Platform captions are small, plainly styled, sometimes off by default, and carry no emphasis. Burned-in captions are part of the edit and are visible to everyone regardless of settings.
Should captions be word by word or full sentences?
Neither extreme. One to three words per beat, changing with the rhythm of speech, reads well and keeps the eye moving. Single words feel frantic on longer clips, and full sentences require reading rather than glancing.
Where should captions sit?
Middle third, roughly at or slightly below centre. Platform interface elements cover the bottom of the frame on every major app, and captions placed there get obscured by the caption text, the username and the buttons.
Do captions change performance?
They change whether a muted viewer can follow the clip at all, which for a talking-head format is close to determining whether it works. It is less a performance tweak than a requirement for the format.
Want your podcast turned into clips that hold?
We cut, caption and distribute short-form for podcasts and founders. Bring one episode and we will show you what comes out of it.
Book a call