A university streams its board meeting live with automatic captions. Latency is fine, accuracy is acceptable, nobody complains.
Two hours later, the recording is posted to the public archive. Same caption file. Same errors. Except the recording is no longer live, and nothing about it is judged as live anymore.
That is the single most common captioning compliance failure in media and education right now. Not a bad live caption. A live caption that was never cleaned up before it became on-demand content.
Understanding why requires understanding that live and post-production captions are not two quality grades of the same thing. They are two different jobs.
Two different jobs, not two tiers of quality
Live captioning is a real-time service. A human or a machine converts speech into readable text within seconds, with no opportunity to review, and delivers it into a signal that is already moving. Success means the viewer follows along without falling behind.
Post-production captioning, sometimes called offline captioning, is asset creation. The audio already exists. There is time to transcribe verbatim, punctuate correctly, identify speakers, add sound cues, time each caption to the frame and position it so it never covers a graphic.
One is judged on keeping up. The other is judged on getting it right. Regulators apply that distinction deliberately, which is where the trouble starts for teams that treat a live file as a finished deliverable.
The comparison at a glance
| Dimension | Live captions | Post-production captions |
| Produced | In real time, during transmission | After the audio exists, before publication |
| Typical method | Stenography, re-speaking, or automatic speech recognition | ASR first pass plus human review and conformance QC |
| Accuracy standard | NER, with 98% the accepted live threshold | Verbatim and error-free, commonly 99%+ measured on WER |
| Latency | 1 to 3 seconds is a common live streaming target | Not applicable, timing is frame-accurate |
| Display style | Roll-up, usually two or three lines | Pop-on, positioned and timed per caption |
| Common formats | 608 or 708 embedded, streaming WebVTT segments | SCC, MCC, STL, SRT, WebVTT, TTML or IMSC |
| Editable after the fact | No, but the output can be cleaned into a VOD file | Yes, before delivery |
| Cost driver | Scheduled hours and captioner availability | Runtime, complexity and turnaround |
| Fails when | Audio is poor, speakers overlap, jargon is unregistered | Reference styles, platform specs and placement are ignored |
How live captioning actually works
Three delivery methods dominate, and they are not interchangeable.
Stenographic captioning, often called CART, uses a trained stenographer on a machine shorthand keyboard. Highest accuracy on difficult live content, highest cost, and constrained by captioner availability. Still the standard for courts, legislatures and high-stakes broadcast.
Re-speaking, or voice writing, has a trained captioner listen and re-speak the content into a speech engine tuned to their voice, correcting as they go. It handles noisy source audio well because the engine hears one clean voice instead of a stadium. This is the method the NER accuracy model was designed to score, because re-speakers condense deliberately.
Automatic captioning sends the programme audio straight to ASR. Fastest to deploy, lowest cost, most exposed to accents, overlapping speech, crowd noise and unregistered proper nouns. Our guide to automatic closed captioning software covers where that first pass is genuinely sufficient.
Whichever method you use, the output has to reach the viewer. In broadcast that means injection into the signal through a caption encoder. In streaming it means segmented caption data travelling with the manifest. Both add latency, which is why the number your vendor quotes should be end to end, not engine-only. We go deeper on live delivery in our guide to live captioning services for broadcast events.
How post-production captioning works
The offline path looks slower on paper and is usually cheaper per finished minute, because the constraint is runtime rather than scheduling a specialist for a fixed window.
A typical pipeline runs ASR for the first-pass transcript, human review for accuracy, speaker labels and sound cues, then caption segmentation against reading rate and line-length rules, then placement so captions clear lower thirds and on-screen text, then conformance checking against the destination platform’s spec, then export in whichever format that platform demands.
That last step catches more teams than the accuracy step. A file can be perfect and still be rejected because the platform wanted IMSC and received SRT, or because reading speed exceeded the spec. Our closed captioning guidelines break the platform-by-platform differences down.
What the rules actually expect from each
This is where the distinction stops being academic.
The FCC holds captions to four standards: accuracy, synchronicity, completeness and placement. For live and near-live programming it evaluates against a “greatest extent possible” test that accounts for the realities of real-time transcription. For prerecorded programming there is no such allowance, because there is no technical excuse for gaps. Our breakdown of FCC caption rules for TV and streaming covers where the obligation sits.
WCAG draws the same line differently. Captions for prerecorded media sit at Level A under success criterion 1.2.2. Captions for live media sit at Level AA under 1.2.4. Any organisation targeting WCAG 2.1 AA owes both.
That target is now a legal deadline for a large group. The DOJ’s ADA Title II rule requires state and local public entities to meet WCAG 2.1 Level AA, and in April 2026 the Department issued an interim final rule extending the compliance dates by a year: April 26, 2027 for entities serving 50,000 or more people, and April 26, 2028 for smaller entities and special districts. Public meetings, lecture capture and archived video all fall inside scope.
The handoff nobody owns
Back to the board meeting.
A live stream is judged as live. The moment that same recording is published on demand, it is prerecorded content and the flexibility disappears. If the live caption file travels with it unchanged, you have published a prerecorded asset against a live standard.
The fix is a defined post-live pass, and it is cheap because you are not starting from zero.
- Capture the live caption output as a working transcript, not a throwaway
- Run a correction pass against the recorded audio for names, numbers and negations
- Re-segment from roll-up into pop-on, timed to the frame
- Check placement against graphics and lower thirds in the recorded version
- Re-export in the format the VOD platform requires
- Replace the file before or at publication, not weeks later
- Assign an owner, because this task sits between two teams and usually falls between them
Set a service level on it. Most operations can turn a clean VOD file inside 24 hours of the event.
Choosing the path for each piece of content
- Live-only, no archive. Live captions, method chosen by risk. Automatic for internal town halls, re-speaking or CART for regulated or public-facing events.
- Live now, on demand later. Live captions plus a mandatory post-live cleanup. This is most sport, news, education and public sector video.
- Prerecorded, published once. Post-production only. There is no reason to accept live-grade accuracy on content that was never live.
- Prerecorded but urgent. Post-production with a rush turnaround. Faster than you expect, and still verbatim.
If you are not sure how your current output scores on either path, our explainer on caption accuracy and word error rate covers how to benchmark it properly.
Common objections, answered
“Automatic captions are good enough for live, so they are good enough for VOD.” Different standard, different scrutiny. Live tolerance exists because correction is impossible in real time. On demand, correction is possible, so it is expected.
“Post-production captioning is too slow for our schedule.” Turnaround is a commercial variable, not a technical ceiling. Rush offline workflows routinely deliver same day.
“We cannot afford stenographers for everything.” Nobody can. That is the point of tiering by risk. Reserve CART for content where an error carries legal or reputational cost, and use re-speaking or automatic capture elsewhere.
How Digital Nirvana handles both paths
Running two captioning workflows through two vendors is how the VOD handoff goes missing. Digital Nirvana builds them as one pipeline.
TranceIQ covers cloud transcription, caption and subtitle generation, translation and conformance checking, with human review inside the workflow rather than bolted on afterwards, and export into the broadcast and streaming formats your destinations expect. Media Enrichment supplies the live captioning capacity and the trained specialists who run the post-live correction pass, apply your style guide and clear the errors no automated score catches.
Because both paths share one platform, the live output becomes the starting point for the on-demand file instead of a dead end. Our customer success stories show what that looks like across broadcast, streaming and education deployments.
Where captioning connects to the rest of your operation
A finished caption file is the cheapest metadata you will ever produce, and the post-production path is where it becomes usable.
Once a transcript is verbatim and frame-accurate, it stops being a compliance artifact and becomes an index. MetadataIQ writes time-coded transcripts into PAM and MAM environments so producers search for a moment instead of scrubbing a file, and archive teams can license clips they can actually find. The same text layer improves recommendations, contextual ad placement and AI summarisation.
The pattern repeats outside media. Learning management workflows run live captions for lectures and clean transcripts for the course archive, which is exactly the two-path model, with students rather than viewers on the receiving end.
Frequently asked questions
What is the difference between live and post-production captions? Live captions are generated in real time during transmission, with no opportunity for review. Post-production captions are created after the audio exists, allowing verbatim accuracy, frame-accurate timing, speaker labels and controlled placement.
Can I reuse my live caption file for the on-demand version? Only after a correction pass. Once content is published on demand it is treated as prerecorded, and prerecorded content is held to a stricter standard than the live original was.
How accurate should live captions be? The widely accepted threshold for live captioning is 98% measured on the NER model, which weights errors by how much they damage comprehension rather than counting every word equally.
Which is more expensive, live or post-production captioning? Live usually costs more per minute because it depends on scheduled specialist availability. Post-production cost scales with runtime, complexity and turnaround speed.
Ready to close the gap between your live and on-demand captions? Book a 15-minute captioning workflow review and we will map where your live output goes after the stream ends.
Conclusion
Live captions and post-production captions are not competing options. They are sequential stages of the same obligation, and the failure point sits between them.
Choose your live method by how much an error costs. Choose your post-production turnaround by when you publish. Then, most importantly, own the handoff, because the moment a live stream becomes an on-demand asset, the standard changes and nobody sends a reminder.
Key takeaways
- Live captioning is a real-time service judged on keeping up. Post-production captioning is asset creation judged on getting it right.
- Live accuracy is best scored with NER at a 98% threshold. Prerecorded work is held to verbatim, error-free output.
- The FCC judges live and near-live content on a greatest extent possible basis. Prerecorded content gets no such allowance.
- WCAG puts prerecorded captions at Level A and live captions at Level AA, so WCAG 2.1 AA requires both.
- ADA Title II compliance dates were extended in April 2026 to April 26, 2027 and April 26, 2028.
- Publishing an uncorrected live caption file as VOD is the most common avoidable compliance failure. Build a post-live pass and give it an owner.