Captioning works well when it is a standing process, video comes in, captions go out on a known timeline, and works far less well when it is a queue of individual requests. The difference is mostly organizational, and most of the real choices come down to how much accuracy a given piece of content needs, and how fast.
Here are the models institutions use, and how they tend to combine them.
The building blocks
Automatic speech recognition, or ASR, is fast, inexpensive, and steadily improving. It runs in the 85 to 95 percent accuracy range for clear audio and general vocabulary, lower for technical terms, accents, crosstalk, or poor room audio. ASR captions are useful immediately, for search, for a lot of students as-is, and as a solid first draft.
Human review and correction is a person cleaning up the ASR output: fixing terms, speaker labels, punctuation, and anything the model simply misheard. This is where captions reach the accuracy a formal accommodation or reusable course content needs. It takes roughly three to six times the runtime of the video for a skilled editor working from a decent ASR draft.
Full human transcription is captioning from scratch, used mainly when the audio is too poor for ASR to produce a usable draft at all. It is slower and more expensive, and worth avoiding wherever you can just improve the source audio instead.
Common models institutions run
ASR only, applied automatically. Every recording gets machine captions on publish, no routing involved. This is solid baseline coverage, appropriate for the sheer volume of day-to-day lecture capture, and honest about its own limits, which is exactly why most institutions pair it with a clear path to request corrected captions.
A supplier service for corrected captions. A commercial captioning supplier handles the human-review step, usually with tiered turnaround. Standard in four to five business days, rush in one to two at a higher rate. It is predictable, it scales cleanly with spend, and it removes the staffing question entirely. Most institutions use one for at least their accommodation-driven and high-priority work.
An in-house captioning team. Staff or trained student workers handle the review step directly. Lower marginal cost at volume, better handling of course-specific terminology, and useful when turnaround needs to be tight and tightly controlled. It requires real management and a proper tool, and its capacity is fixed, so most in-house teams end up paired with a supplier for overflow.
Hybrid, sorted by content type. This is where most institutions land: ASR automatically on everything; in-house or supplier human review for accommodation requests, high-enrollment courses, reusable lecture content, and anything an instructor flags directly; full transcription reserved only for salvage cases.
Sizing it to real demand, not total volume
The volume that needs human-accuracy captions is much smaller than total video volume. Institutions size their human-review capacity to accommodation requests, which come from disability services and are the one category with a hard turnaround expectation attached; to high-enrollment courses, where the number of students served per hour of editing is highest; to reusable content, lecture material shown again across terms, which earns the full review exactly once; and to instructor-flagged content, handled through a simple “request corrected captions” button on each recording. Everything else stays on ASR, with the correction path sitting there if a need surfaces.
Making it a process, not a queue
A few practices make the real difference. Publish turnaround commitments. ASR captions within an hour of publish, corrected captions within a stated number of business days. So faculty and disability services can plan around known numbers instead of guessing. Wire captioning directly into the capture pipeline so it happens automatically on publish rather than as a separate upload someone has to remember. Improve source audio, because better room microphones raise ASR accuracy across the board and shrink the human-review load at the same time — audio investment pays off here just as much as it does in the live experience. Keep a course glossary for in-house editors, listing the recurring technical terms in a program, so corrections stay consistent and get faster over time. And track backlog and turnaround as standing metrics, the same way a help desk tracks its tickets.
A reasonable starting point
Institutions building this up from a request-based system usually sequence it the same way: turn on automatic ASR for all recordings first, then stand up a modest human-review capacity, in-house or supplier, sized to accommodation requests plus the top-enrollment courses, publish the turnaround commitments, and add the instructor-flag button last. That covers the great majority of real need, and turns captioning from a recurring scramble into an actual service.
Emily Connelly is an instructional multimedia technologist at Wrightmann Education Technologists. Part of the Accessibility and Compliance series.
References
- W3C. (2018). Web Content Accessibility Guidelines (WCAG) 2.1. Time-based media. https://www.w3.org/TR/WCAG21/
- U.S. Department of Justice, Civil Rights Division. (2024). Fact Sheet: New Rule on the Accessibility of Web Content and Mobile Apps. https://www.ada.gov/resources/2024-03-08-web-rule/