How to Add Subtitles and Captions to Any Video (Free Tools That Work)
Why captions boost accessibility and retention, SRT vs VTT, the best free auto-transcription tools, burned-in vs closed captions, translation, and QC checklists.
Originally published on Video Editing by Omar Haddad. Read on the original site
Captions used to be an afterthought — a compliance checkbox for broadcasters and a courtesy for the few. Today they are core infrastructure for every video you publish. A huge share of social video plays with the sound off, viewers increasingly leave captions on by default even with sound available, and platforms index caption text for search and recommendations. Meanwhile, the tooling has been transformed: speech recognition that once cost dollars per minute is now free, fast, and accurate enough that the excuse for uncaptioned video has evaporated. This guide covers the whole workflow — formats, the best free tools, styling, translation, and the quality checks that separate professional captions from machine sludge.
Why Captions Matter More Than You Think
Two arguments, and both are strong.
Accessibility. Hundreds of millions of people worldwide live with disabling hearing loss, and captions are their primary access to your content. Beyond deaf and hard-of-hearing viewers, captions serve people watching in noisy environments, non-native speakers who read better than they parse spoken language, and viewers with auditory processing differences. In many jurisdictions and for many institutional clients, captions are also a legal or contractual requirement, not a nicety.
Retention. Creators who A/B test captions almost universally keep them, because captioned videos hold viewers longer. The mechanics are intuitive: a viewer scrolling with sound off gives your video two seconds — captions are the only way your hook lands in that window. Captions also reinforce comprehension (viewers who read and hear a point retain it better) and give the eye something to track during slower moments. Add the discoverability effect — platforms and search engines index caption text — and captions are one of the highest-return five-minute tasks in publishing.
One vocabulary note before we go further. Subtitles traditionally assume the viewer can hear and translate or transcribe dialogue only; captions (specifically closed captions) also describe non-speech audio — [door slams], [tense music] — for viewers who cannot hear it. In casual use the words blur together, but if you are captioning for accessibility, include those non-speech cues.
SRT and VTT: The Two Files You Need to Know
Nearly every captioning workflow ends in one of two plain-text formats.
SRT (SubRip) is the ancient, universal standard. Numbered blocks, a timecode range, and text:
1
00:00:01,000 --> 00:00:03,400
Welcome back to the channel.
2
00:00:03,600 --> 00:00:06,200
Today we're fixing your caption workflow.
VTT (WebVTT) is the web-native evolution — required by HTML5 video players and preferred by some platforms. It looks nearly identical but starts with a WEBVTT header, uses periods instead of commas in timestamps, and supports styling and positioning cues that SRT cannot express.
| Feature | SRT | VTT |
|---|---|---|
| Platform support | Nearly universal (YouTube, editors, players) | HTML5 players, YouTube, modern platforms |
| Styling/positioning | None (plain text) | Yes (CSS-style cues, positioning) |
| Timestamp format | 00:00:01,000 (comma) | 00:00:01.000 (period) |
| Best use | Uploads, archives, maximum compatibility | Web embeds, styled closed captions |
The practical advice: work in SRT by default, because everything accepts it, and convert to VTT when a web player demands it. Conversion is trivial — the formats are so close that free converters (or a careful find-and-replace) handle it in seconds. Whatever you do, keep the caption file alongside your master video file forever; it is your transcript, your translation source, and your re-upload insurance.
The Free Tools That Actually Work
Auto-transcription has become a solved problem for clear speech in major languages, and the free options are genuinely good. Here is the current landscape.
YouTube Studio auto-captions everything you upload, in dozens of languages. Accuracy on clear speech is solid, and the built-in caption editor lets you fix errors and adjust timing right in the browser — then download the result as SRT or VTT. A classic free workflow: upload a video unlisted purely to harvest and polish its auto-captions, download the file, and use it anywhere.
CapCut generates auto-captions in one tap, on both mobile and desktop, and is the default choice for short-form burned-in captions. Its templates handle the styling (including word-by-word highlight animation), and its accuracy on conversational speech is among the best of the consumer tools. The free tier covers caption generation; some animated presets sit behind the Pro tier.
Whisper-based tools are the power option. OpenAI's Whisper model was released openly, and an ecosystem of free tools wraps it: desktop apps and front-ends that run entirely on your own computer, no upload, no account, no length limits. Whisper's accuracy — especially on accented speech, technical vocabulary, and noisy audio — beats most built-in platform transcribers, and it outputs SRT and VTT directly. If you caption a lot of long-form content, a local Whisper tool is worth setting up once; your only cost is compute time.
Descript takes a different angle: it transcribes your video into an editable document, and editing the text edits the video. The free tier includes limited transcription hours per month — enough for light use — and its correction interface (play, click a word, retype it) is the fastest way to fix a transcript that exists. Export supports SRT.
Editors' built-in transcription rounds out the field: DaVinci Resolve, Premiere Pro, and Final Cut all now include speech-to-text with caption track generation, so if you are already editing in one of them, captioning happens without leaving the app. Resolve's is included in the free version.
Burned-In vs. Closed Captions: Choose Deliberately
There are two fundamentally different ways to deliver captions, and choosing wrong causes real problems.
Burned-in (open) captions are rendered permanently into the video's pixels. Viewers cannot turn them off, and you control exactly how they look — font, animation, position, the karaoke-style word highlighting that dominates short-form. Burn in your captions when you publish to TikTok, Reels, and Shorts, where styled captions are part of the content's visual identity and where platform caption rendering is inconsistent.
Closed captions live in a separate file or track (your SRT/VTT) that the player renders on demand. Viewers can toggle them, resize them, and — critically — screen readers, search indexing, and automatic translation can access the text. Use closed captions for YouTube long-form, web embeds, courses, and anything with accessibility obligations.
The professional answer for most creators is both: burn styled captions into short-form cuts, and upload a clean SRT alongside long-form versions. Two cautions on burned-in captions specifically. First, they are irreversible — a typo means re-exporting the video, so proofread before the render. Second, keep them inside the platform's safe area: TikTok and Reels overlay UI on the bottom and right of the frame, so captions belong in the lower-middle to middle of the frame, not at the very bottom edge where a username will sit on top of them.
Styling rules that hold across platforms: a heavy sans-serif font, high contrast (white text with a black outline or background block survives every background), no more than two lines at a time, and a reading pace of roughly 15–20 characters per second. If a caption flashes by faster than you can read it aloud, split it.
Translating Subtitles
Once you have an accurate caption file in the original language, translation multiplies your audience for almost no cost — the timing work is already done, and machine translation operates on the text alone.
The realistic quality ladder:
- Platform auto-translation. YouTube can auto-translate your captions into the viewer's language on the fly. Quality is serviceable and the effort is zero — make sure you have uploaded a corrected caption file, because auto-translating an error-filled auto-transcript compounds mistakes badly.
- Machine-translating your SRT. Free translation tools handle SRT text well, and dedicated free subtitle translators preserve the timecode structure while translating only the dialogue lines. This gives you an uploadable foreign-language subtitle track for each target market.
- Native-speaker review. For your top one or two languages by audience size, pay or ask a native speaker to review the machine output. Idioms, humor, and culturally specific references are where machine translation still faceplants, and a one-hour review catches nearly all of it.
Practical notes: translated text runs longer than English in most European languages (German famously so), so check that translated lines still fit in two lines and remain readable at your timing. Keep line breaks at phrase boundaries after translation, not just wherever the character count happened to land. And upload each language as its own caption track rather than burning any translation in — closed caption tracks let one video serve every market.
Quality Control: The Ten-Minute Pass That Saves Your Reputation
Auto-transcription gets you 90–97% of the way there, and the remaining few percent is where the embarrassing failures live. A repeatable QC pass:
- Watch with sound off. You will experience the captions the way a deaf viewer or muted scroller does, and timing problems become obvious.
- Fix proper nouns first. Names, brands, and technical terms are where every engine fails. Search the transcript for your own product name — it is wrong more often than you would hope.
- Check number and homophone traps. "Their/there," "to/too," spoken numbers ("twenty twenty-six" vs "2026"), and units are classic silent errors.
- Verify sync at three points — beginning, middle, and end. Drift usually accumulates, so if the end is in sync, the middle almost certainly is.
- Confirm speaker changes are clear. In multi-speaker content, add speaker labels or dashes when it is not visually obvious who is talking.
- Add non-speech cues if this is an accessibility caption track: [laughter], [phone buzzes], [music fades].
- Read the captions as a document. Export the text and skim it top to bottom; errors that hide at playback speed jump out on the page.
Punctuation deserves a special mention because auto-transcribers are still mediocre at it, and captions without punctuation are exhausting to read. Sentence-ending periods and question marks matter more than commas; fix those at minimum.
A Workflow to Steal
Pulling it all together, here is a caption pipeline that costs nothing and scales from a single Short to a weekly long-form channel: transcribe with the best free tool for your context (CapCut for short-form, Whisper or YouTube Studio for long-form), correct the transcript with the QC pass above, export SRT as your master caption file, burn styled captions into vertical cuts, upload the SRT as closed captions everywhere else, and machine-translate the corrected SRT for your biggest non-native audiences. The first video takes half an hour to set up; every one after takes ten minutes. Few things in video production return this much audience for this little effort.
FAQ
Are auto-generated captions good enough to publish without editing?
No. Modern engines are impressively accurate on clear speech, but they still miss proper nouns, punctuation, and homophones — and an uncorrected error in a caption sits on screen in writing, which viewers judge more harshly than a spoken stumble. Budget ten minutes of correction per video; it is the highest-value part of the whole workflow.
What's the difference between subtitles and closed captions, practically?
Subtitles carry dialogue text and assume the viewer can hear everything else; closed captions also describe meaningful non-speech sound for viewers who cannot hear it. If accessibility is a goal (and it should be), include the non-speech cues — it is a few extra minutes of work.
Should captions be at the top or bottom of a vertical video?
Neither extreme. Platform UI covers the very bottom, and the top competes with your hook text. The sweet spot for TikTok, Reels, and Shorts is the lower-middle of the frame — roughly the 60–75% vertical mark — where captions stay clear of usernames, buttons, and the viewer's focus on faces.
How do I caption a live stream?
Live captioning is a different problem — you need real-time speech recognition. YouTube offers automatic live captions on streams, and OBS integrates with free live-captioning plugins. Quality is below recorded-and-corrected captions, so for important streams, publish a corrected VOD with proper captions afterward.
Originally published on Video Editing by Omar Haddad. Read on the original site