Industry Insights · May 23, 2023

New Progress in Speech Translation: From 'Translator Devices' to 'Live Captions'

Audience member holding a phone with a live-captioning app

In 2023, speech translation technology took a new leap in usability. Combining automatic speech recognition (ASR) with large language models improved both "hearing correctly" and "translating accurately": recognition rates under accents, fast speech and background noise rose, and output fluency gained from LLMs' language ability. Scenarios like live meeting captions, cross-language voice calls and video speech translation are moving from keynote demos to daily utility.

Video-conferencing platforms are the main carrier of this progress: mainstream platforms' live-caption features now support multilingual translation, letting meeting participants follow along with captions in their native language. While quality is still far from replacing professional interpreters, as a "comprehension aid" it meaningfully lowers the barrier to joining cross-language meetings.

The industry's positioning of speech translation is also growing more rational: it solves the problem of "whether communication exists", not "how good it is". In high-risk scenarios — business negotiation, medical consultation, court proceedings — professional interpreters remain irreplaceable: live captions can help you grasp the gist, but no one wants to bear the consequences of a captioning error.

For interpreters, progress in speech translation is less a threat than a refinement of the division of labor: machines cover the long tail, humans hold the high line — a pattern that grew still clearer in 2023.

Let's talk about your language needs

Tell us about your project — we'll reply with a quote and delivery plan within one business day.