← Back to blog
NewsAbout 6 min read

Tavus Griffin Passed as Human 48 Percent of the Time, and the Company Is Not Shipping It

Published Oct 7, 2026
Tavus Griffin Passed as Human 48 Percent of the Time, and the Company Is Not Shipping It

Tavus released Griffin on October 1, describing it as the first Human Interaction Model: a full-duplex video-to-video system that listens, reacts, and speaks in real time. In a live video call study, 26 of 54 participants believed Griffin-Lite was a human after a minute of conversation. The previous generation of the system managed 1 out of 41, or 2.4 percent.

That is roughly a twentyfold improvement in convincing people that a synthetic participant is real. Tavus has stated plainly that it will not open Griffin to the public until safety mechanisms are ready. Given what the model does and what the failure modes look like, that decision deserves more attention than the demo reel.

The technical picture

Griffin runs at 720p and 25 frames per second, with audio-to-video latency of about 0.43 seconds. On the perception track of NVIDIA's VideoFDB benchmark it ranks first at 3.73, against a human baseline of 4.20. It ranks second on lip-sync accuracy and is slower to respond than a person.

The access threshold is the part that matters. A single photograph plus ten seconds of audio is enough to generate a working digital double. Tavus's API has roughly 150,000 developers and businesses registered.

Put those two facts next to each other and the risk profile is not hypothetical. Surfshark data puts deepfake fraud losses at $1.1 billion in 2025, three times the 2024 figure. In the 2024 Arup case in Hong Kong, an employee joined a video conference where every other participant was an AI fabrication, and the company lost $25 million.

Griffin does not make deepfakes possible; that has been achievable for years. It makes them interactive. A pre-recorded fake can answer only the questions the creator anticipated. A real-time system answers whatever it is asked, in a conversation the other party believes they are having with a colleague.

What changes when the fake can improvise

The Arup case worked because the fraudsters had prepared a scenario and the target accepted it. Real-time generation removes the preparation constraint. A caller can now improvise through an unexpected question (asking about a project detail, referencing a shared meeting, reacting to a joke) and generate a plausible response in under half a second.

That is where the 0.43 second latency figure stops being a spec sheet entry. Conversational turn-taking in humans runs around 200 milliseconds of gap. At 430 milliseconds, Griffin is noticeably slow but not outside the range of someone on a poor connection or in a noisy room. The gap is maskable by context, which is exactly what makes it dangerous: users attribute it to network conditions, not to something being wrong.

Tavus's decision not to ship reflects an accurate read of this. The company is between two positions. The technology is demonstrably capable of fooling a majority of participants in a controlled study. The countermeasures (provenance marking, liveness detection, verified identity in calls) are not in place at the scale the technology would need.

The benchmark has a human baseline, and that is the point

Worth noting that VideoFDB reports a human score of 4.20 against Griffin's 3.73. The model is ranked first among systems and still below people. That gap is the current state of the art, and it is shrinking at a rate that makes "below human" a declining comfort.

The more useful number is participant belief rate over time. The study measured belief after one minute. Most detection strategies depend on the observer noticing something over a longer window, a repeated gesture, a response that does not fit the conversation history, a question dodged twice. A system that survives the first minute and degrades over five is a different threat than one that holds up for an hour, and the published figure covers only the first.

What organizations should be doing now

The defensive advice is unglamorous and mostly organizational.

Text-based verification of any financial instruction is no longer sufficient, and neither is voice. A callback to a known number is better than a video call to an unknown participant, because the channel is verified even when the person on it might not be. Shared secrets work until they leak, which they do.

The stronger measure is procedural: high-value instructions should require a second, independent channel, and that second channel has to be one the requester cannot initiate. This is standard practice in treasury operations for exactly this reason, and it scales to a world where any video participant can be synthetic.

For platform operators, the pressure is on liveness and provenance. C2PA manifests and SynthID watermarks attach provenance to generated media, and neither survives a re-encode reliably. Real-time video has no manifest at all, which means the verification has to happen on the receiving end, during the call, on the basis of behavioral signals that are exactly what Griffin is trained to produce plausibly.

There is an unpleasant circularity here worth naming. The behavioral signals people use to judge whether a video participant is real (natural pauses, micro-expressions, appropriate reactions to unexpected content) are the same signals a real-time model is optimized to reproduce. Detection based on those signals puts defenders in a permanent trailing position, always checking for the last generation's tells after the current one has learned to avoid them.

The signals that are harder to fake are the ones outside the model's control: whether the identity claim is verifiable through an independent channel, whether the account has a history, whether the same person can be reached a second time by a different route. Identity infrastructure rather than behavioral detection.

The disclosure standard is the real fight

Tavus's restraint is a company decision, and company decisions are reversible. The durable question is what the default should be when a system can pass as human in a majority of short interactions.

The comparison to draw is the EU AI Act and the synthetic content labeling regimes now taking effect globally. Those rules were designed around media that can be marked after generation, with metadata and visible labels. Real-time synthetic video is not that. There is nothing to label, because the output is ephemeral and conversational. The disclosure has to be a property of the system making the call, a registered, verifiable identity for the automated participant, which the receiving party can check.

That is a solvable engineering problem, and nobody has shipped it. Until someone does, the honest state of affairs is that a company has built something capable of convincing most people that a machine is a person on a video call, published the data showing it works, and declined to release it. That is a better outcome than the alternative, and it is a decision that holds only as long as one company wants it to.

Related articles