Tavus introduces Griffin, a model it says passes the video Turing test

Tavus launched Griffin on October 1, 2026, a real-time video conversation model that fooled 48% of study participants into thinking it was human, versus a 2% pass rate for prior systems.

Tavus introduces Griffin, a model it says passes the video Turing test

What Griffin is

Tavus introduced Griffin on October 1, 2026, describing it as the first Human Interaction Model (HIM), a system designed to understand and generate face-to-face real-time conversation. In a live study, 48% of participants who talked with Griffin believed they had spoken with a real human, compared with a maximum 2% pass rate for previous systems, including Tavus's own conversational video interface.

The model combines perception, conversational decision-making, and expressive video generation into one video-to-video system. Griffin-Lite, a research preview, is available to a select group of early testers, with a wider release planned after safety evaluations.

How it works

Griffin differs from turn-based systems that wait for a speaker to finish before responding. Instead, it makes conversational decisions at sub-second intervals, deciding whether to speak, nod, backchannel, or stay silent while continuously processing audio and video. It can interrupt, yield, or adjust mid-sentence based on what it sees and hears.

The system has two main parts. A continuous conversational modeling engine perceives incoming audio and video and produces control signals for what to say and how to say it. An audiovisual generation engine converts those signals into speech and video in real time, using a fast autoregressive diffusion transformer for voice and a few-step autoregressive generator for 720p video in 320 ms chunks.

Hassaan Raza, co-founder and CEO, framed the goal in terms of reducing friction in human-machine interaction: "Humans are evolutionarily designed to communicate face to face. We speak as much through our words as we do our expressions, tone, gestures and timing." The aim, he said, is for computing to "become invisible."

Benchmarks and the video Turing test

On NVIDIA's VideoFDB benchmark, which evaluates full-duplex audiovisual conversation, Griffin-Lite scored highest on both the generation and perception tracks against published baselines. On generation, it scored 3.83, compared with 3.92 for the human reference and 2.80 for the next-best system. On perception, it scored 3.73, ahead of the strongest reported baseline at 3.44 and the human reference at 4.20.

In the face-to-face study, 54 participants had a one-minute video call with Griffin-Lite without being told their partner was an AI model. Of those, 26 believed they had spoken with a real person. Participants rated the model 5.4 out of 7 for seeming natural, 5.6 for seeming trustworthy, and 5.8 for whether they would enjoy talking with it again.

The video generator also outperformed four published streaming diffusion models on latency, averaging 0.43 seconds from audio input to visible response on H100s, half the time of the next fastest method. It ranked first on perceptual quality (DOVER), distribution fidelity (FID), and the Talking Head Evaluation framework (THEval).

Safety approach

Tavus said Griffin-Lite will not be available to customers yet. The company cited the model's ability to pass the Turing test as a reason for caution, noting that the same properties that make it a natural interface also allow it to deceive people into thinking it is not AI. The company is working on disclosure features and collaborating with organizations focused on AI safety before a broader release.

Why this matters for customer support, sales, and marketing teams

For professionals in customer-facing roles, Griffin signals where conversational AI is heading: systems that read tone, pause, and visual context rather than just transcribed words. A support agent evaluating AI tools should expect vendors to start claiming full-duplex, video-based interaction as a differentiator. Sales and marketing teams testing AI avatars for demos or training will need to weigh the disclosure and trust implications of models that people cannot reliably distinguish from humans.

For teams building or buying AI-powered communication tools, the benchmark results matter. The gap between Griffin's VideoFDB scores and existing commercial systems suggests that real-time audiovisual models may soon reset expectations for what counts as natural interaction. Professionals evaluating Generative AI Courses or AI Video Creation Courses should pay attention to how quickly these capabilities move from research previews to production tools.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)