Subtitles look like a transcript and are not one. A good subtitle track is edited, condensed, timed to the frame and constrained by reading speed. When subtitles feel effortless, considerable craft went into them. When they feel wrong, it is usually because one of those constraints was ignored.
Written September 2026. Standards vary by broadcaster and platform, though the principles below are near-universal.

Subtitles, captions and SDH
Three terms get used interchangeably and mean different things.
- Subtitles render dialogue, usually translated, and assume the viewer can hear. They do not describe sound.
- Closed captions are intended for deaf and hard of hearing viewers in the original language, and include speaker identification and meaningful non-speech sound.
- SDH, subtitles for the deaf and hard of hearing, is the streaming-era hybrid: subtitle formatting with caption content. It exists because streaming platforms needed something that worked across delivery systems that did not support traditional broadcast captions.
This is why selecting English subtitles on a film sometimes gives you door creaks and sometimes does not. They are different tracks made for different audiences.
How subtitles are produced
- Transcription or speech recognition produces a raw text pass. Automatic recognition has become good enough to use as a starting point, and not good enough to publish.
- Spotting sets the in and out times for each subtitle against the picture. Good spotting respects shot changes, because a subtitle that straddles a cut forces the eye to re-read.
- Condensing and editing reduces the text to something readable in the time available. This is the skilled part and the reason subtitles are not transcripts.
- Translation, where applicable, works within those same constraints, which is why subtitle translation is harder than translating prose.
- Quality checking against reading speed, line length, positioning and any burned-in on-screen text that must not be obscured.
The rules that govern the text
The constraints look arbitrary and are grounded in reading research. Reading speed is typically capped around 17 characters per second for adults and lower for children’s content. Subtitles are limited to two lines, occasionally three. Line breaks are placed at grammatical boundaries rather than wherever the line runs out. A minimum duration prevents flashing, and a small gap between subtitles signals that the text has changed.
Condensing to fit those limits is why subtitles sometimes differ from the audio. It is not sloppiness. A faithful transcript of fast dialogue would be unreadable, so the subtitler preserves meaning and drops what the viewer can infer.
Why some subtitle tracks are noticeably worse
- They were machine generated and not edited. Recognisable by missing punctuation, no speaker identification, and no condensing.
- They were translated from another translation. Pivot translation through English is common and compounds errors.
- The timing ignores shot changes, which is fatiguing in a way viewers feel without being able to name.
- They were made for a different cut of the film, so drift accumulates across the running time.
- Reading speed was ignored, producing walls of text that vanish before they can be read.
Why more people are using them
Subtitle use has risen sharply among viewers with no hearing loss, and the causes are practical rather than mysterious. Modern productions favour naturalistic, quietly delivered dialogue. Audio is mixed for cinema and then played on flat television speakers and phones. A great deal of viewing happens in noisy places or with the volume down. Subtitles compensate for all of it.
If dialogue is hard to follow at home, subtitles are the reliable fix, though speaker placement and a centre channel help considerably too. Our home cinema setup guide covers the audio side of the same problem, and streaming live sports covers a context where commentary and crowd noise compete constantly.
Styling and where they appear
Placement is part of the craft. Subtitles normally sit centred at the bottom of the frame and move when they would obscure something that matters, such as burned in on-screen text, a face at the bottom of the shot, or a broadcaster logo. Some productions position them beneath the relevant speaker, which helps in scenes with several people talking.
Most platforms now let viewers change size, font, colour and background opacity, and it is worth spending two minutes on those settings rather than tolerating the default. A solid background box costs a little of the picture and dramatically improves readability over bright scenes, which is the single change most people notice immediately.
If you are making your own
For anyone captioning their own video, the practical advice is short. Use automatic recognition for the first pass and then edit every line, because the errors cluster on names and technical terms, which are exactly the words carrying the meaning. Keep to two lines. Break lines at natural grammatical points. Check the result at full speed rather than reading the file, since timing problems are invisible in a text editor and obvious on screen.
Why translated subtitles differ from dubbing
When a film offers both subtitles and a dub, the two scripts are usually not the same, and viewers who switch between them notice the discrepancy. Dubbing is written to match lip movement and timing, which forces different word choices. Subtitles are written to be read at speed, which forces condensing. Both are legitimate translations of the same dialogue solving different problems, and neither is the transcript. It is the most common reason people conclude that one version is inaccurate when both are doing their job.
Common questions
What is the difference between subtitles and captions? Subtitles render dialogue and assume you can hear. Captions are written for deaf and hard of hearing viewers and include speaker identification and meaningful non-speech sounds.
Why do subtitles not match the spoken words? Because they are condensed to a readable speed, typically around 17 characters per second. A literal transcript of fast dialogue could not be read in the time available, so subtitlers preserve meaning over wording.
What does SDH mean? Subtitles for the deaf and hard of hearing: subtitle formatting combined with caption content, created because streaming platforms needed a format that worked across different delivery systems.
Why do so many people watch with subtitles on? Naturalistic dialogue delivery, cinema audio mixes played through flat TV speakers, and viewing in noisy environments or at low volume. It is a response to how content is made and consumed rather than a hearing issue.
Formats and why files sometimes fail
Subtitle files come in a handful of formats, and mismatches cause most of the problems people hit with downloaded tracks. SRT is the simplest and most widely supported, carrying text and timings only. WebVTT adds positioning and styling for web players. Broadcast formats carry considerably more. If a subtitle file loads but appears unstyled or misplaced, it is usually a format the player only partly supports. If timings drift steadily, the file was made for a different frame rate or a different cut, which no amount of player configuration will fix.
Sources and further reading
Where the figures and rules above come from, so you can check them:
- Subtitle standards including timing and reading speed: BBC Subtitle Guidelines
- Accessibility requirements for captioning: W3C Web Accessibility Initiative
Join the discussion