Test your website →

Blog

Captions and Transcripts: Reaching the Muted Majority

Captions were built for deaf viewers, but most of the people who need them today are simply watching with the sound off.

Watch someone scroll through their phone on a train, in a waiting room, or in bed next to a sleeping partner. Video after video plays, and the sound stays off. This is how a huge share of video gets watched now: silently, in places where audio would be rude or impossible. Which means that when you publish a video without captions, you are not merely failing deaf viewers. You are publishing content that most of your audience, in most of the situations where they encounter it, cannot understand.

I find this framing useful because it dissolves a mental category that does a lot of damage. The category is “accessibility features,” imagined as a special ramp bolted onto the side of the building for a small group of people. Captions started that way, as an accommodation for deaf and hard-of-hearing viewers, and that origin still matters; for those viewers captions are not a convenience but the difference between content existing and not existing. But the accommodation turned out to be the better default for nearly everyone. This happens over and over. Curb cuts were for wheelchairs and got taken over by strollers and suitcases. Captions were for deafness and got taken over by open-plan offices and commutes.

A caption, to be precise, is text synchronized with the video, showing the speech and the meaningful sounds as they happen. A transcript is the whole thing as a document you can read at your own pace. They solve different problems and you want both. Captions serve the person watching the video. Transcripts serve the person who doesn’t want to watch a video at all, which is a larger group than video producers like to admit. If I’m evaluating your product and your explanation lives only in a seven-minute video, you are asking me to spend seven minutes finding out whether you were worth thirty seconds. A transcript lets me skim, and skimming is how busy people say yes.

Transcripts do something else that captions can’t: they exist as text on a page, which means search engines can read them. A video, to a crawler, is nearly opaque. It sees a file and whatever title and description you attached. Everything actually said in the video, every question answered, every term of art your customers search for, is invisible. Publish the transcript and all of that becomes indexable. The best explanation of your product, the one your founder gave on camera in an unguarded moment, finally gets a chance to be found. People sometimes ask what content they should write for search, while sitting on hours of recorded talks and demos that contain better answers than anything they’d write cold.

The objection used to be cost, and it used to be valid. Human transcription was slow and expensive enough that skipping it was a defensible business decision. That excuse is gone. Automatic speech recognition is now good enough that the workflow is generate first, then correct, and the correction pass for a ten-minute video is a modest chore, not a project. It’s still a necessary pass. Raw machine captions mangle product names, homophones, and anyone with an accent, and captions with your product’s name spelled three different ways read as carelessness. But editing text is cheap. The economics have flipped, and habits haven’t caught up.

One detail worth getting right, because it separates adequate captions from good ones: captions are not only the words. When something meaningful happens in sound, the caption should say so. A notification chime that the presenter reacts to, music that sets a tone, an off-screen question that prompts the answer. A viewer relying on captions should never watch the presenter respond to something that, as far as the captions are concerned, never happened. The test is simple: could someone watching on mute follow not just the words but the moment.

The accessibility standards ask for exactly what I’ve described, captions for recorded video and text alternatives for audio, and if formal compliance matters to your business, this is one of the clearest requirements to satisfy. But I think compliance is the least interesting reason to do it. The interesting reason is that your video content is probably the most expensive content you make, per minute, and without captions and a transcript it reaches only the fraction of your audience sitting somewhere quiet with the sound on and time to spare. When GazeSite audits a page, embedded media without text alternatives is one of the things it flags, and I notice the same pattern repeatedly: the more effort went into the video, the less went into making it reachable. That’s backwards. The muted majority is not going to turn the sound on for you. Meet them where they are, which is silent, skimming, and ready to leave.

More articles

← All posts