Alibaba releases real-time interpretation model that cuts average lag to 2.3 seconds across 60 languages

Alibaba's Qwen3.8-LiveTranslate cuts simultaneous translation lag to 2.3 seconds across 60 languages, an 18% drop from 2.8 seconds. Speech-to-speech translation costs roughly $1.54 per hour via API.

Alibaba releases real-time interpretation model that cuts average lag to 2.3 seconds across 60 languages

Alibaba's Qwen team released Qwen3.8-LiveTranslate on September 19, 2026, a real-time simultaneous interpretation model that cuts average translation lag to 2.3 seconds across 60 languages. For professionals in customer-facing roles, healthcare consultations, or global events, that 18% reduction in delay - down from 2.8 seconds - can mean the difference between a conversation that flows and one that stalls.

The model listens to live speech, optionally processes video frames, and returns translated text and speech while the speaker is still talking. It is available now as a hosted API on Alibaba Cloud Model Studio and QwenCloud under the model ID qwen3.8-livetranslate-flash-realtime over WebSocket.

What changed under the hood

Simultaneous interpretation forces a tradeoff. Wait longer and the model has more context for accuracy. Speak sooner and the listener gets lower latency. Qwen3.8-LiveTranslate rebuilds this loop with what the team calls an Interleave architecture. The latency metric used is LAAL, or Length-Adaptive Average Lagging, which measures how far the translation trails the source speech on average without rewarding systems that over-generate output. The drop from 2.8 seconds to 2.3 seconds represents roughly an 18% cut.

The model builds on the Qwen-Omni stack, drawing on large-scale multimodal data, cross-language and cross-modal alignment, and visual enhancement. A Flash variant also supports offline audio and video translation.

Three new capabilities that change live translation

The release adds features that address real-world friction points in multilingual settings. First, real-time speaker diarization lets the model distinguish speakers in multi-party conversations and preserve each voice through more stable voice cloning. The API exposes cloning modes, including an always mode that re-clones before each response for sessions with multiple speakers.

Second, a synchronized bilingual display streams source transcription and translation as separate events, so both appear on screen together. Third, long-context disambiguation uses conversation history to keep names and terminology consistent. A name introduced early in a meeting stays correct throughout the translation.

Languages, inputs, and what it costs

The model understands 60 languages and can speak 29 of them, returning both audio and text. The remaining 31 languages return text only. Speech output covers Chinese, English, Arabic, German, French, Spanish, Japanese, Korean, Hindi, and others. Inputs are audio and optional images. Visual cues such as lip movements, gestures, and on-screen text help the model handle noisy rooms and ambiguous words. The documentation recommends sending no more than 2 images per second. Teams can also set hotwords - source-to-target term mappings - with a suggested cap of 1,000.

Singapore list pricing per 1M tokens: audio input at $7.50, image input at $0.55, text output at $20, and audio output at $30. Beijing pricing runs lower. Audio input consumes 7 tokens per second and audio output consumes 12.5 tokens per second. One hour of speech in and speech out costs about $1.54 in Singapore, before text and image tokens. The context window is 53,248 tokens, with 49,152 for input and 4,096 for output. Default rate limits sit at 10 requests and 100,000 tokens per minute.

Developers connect through the WebSocket Realtime API. The default turn detection type is speaker_detection. Clients stream audio continuously and receive server-generated responses. Default audio is 16 kHz PCM in and 24 kHz PCM out. The default voice is Tina. Setting session.output_modalities to text only, or text and audio, controls what the client receives. Always send session.finish before closing the connection, or the final segment is lost. Model Studio lists function calling, structured outputs, batch inference, and fine-tuning as unsupported.

Why this matters for customer support, healthcare, and events teams

For support agents handling multilingual calls, a 2.3-second lag with speaker diarization means fewer interruptions and less confusion about who said what. Healthcare interpreters can rely on visual cues - a patient's gesture or a doctor pointing at a chart - to resolve ambiguous terms in real time. Event organizers running live translated panels get synchronized bilingual captions and consistent name handling across long sessions, without post-production delay.

Professionals who want to build these capabilities into their workflows can explore AI Translation Courses or broader AI Engineering Courses to understand the underlying real-time architectures. The model's API-only access means integration requires development work, but the pricing - roughly $1.54 per hour of speech - puts live multilingual conversation within reach for mid-size teams.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)