In Tavus's own study, 120 people watched a 90-second video call and 45% decided the person on screen was real. It was an AI. A separate 120 watched two real humans talk, and 28% decided one of them was the machine.
The company gave me early access to the smaller model it tested, Griffin Lite. The 28% is the number I keep coming back to. More than one in four viewers doubted two real people.
My read is that looking human on video has become a timing problem. People leave gaps of roughly 200 milliseconds between turns, the figure usually cited from a 2009 study of ten languages. Turn-based voice AI waits for silence before it answers, and that pause is often what gives it away.
The company describes this model as full duplex, meaning it watches and listens while it talks. It says the model can be interrupted, cut in on your rambling when asked, and notice someone walking into frame behind you.
NVIDIA researchers built a benchmark for this from 237 clips of real two-person video calls. Real humans score 3.92 out of 5 on its generation track. The company reports 3.83 for its model, against 2.80 for the best system in the published paper. It also reports responses in about 1.9 seconds, where that system took 2.8.
Humans answer in 0.9 seconds on the same test. Nearly a full second still separates the model from a person, and 45% is still under half.
Language models learned what to say. Face-to-face AI is learning when to say it.
Introducing Griffin, the first model to pass the video Turing test.
48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video.
It’s the first Human Interaction Model (HIM).