In a one-minute video call, 26 of 54 people thought the person on the other end was human. The "person" was Griffin-Lite, a version of a new model from the AI video company Tavus. In its October 1, 2026 announcement, Tavus called the result a pass on a "video Turing test." The classic Turing test asks whether a machine can pass for a person in conversation. Tavus's version adds a live face and voice.
Peter Diamandis, founder of XPRIZE, brought the claim to the panel of Moonshots with Peter Diamandis in an episode recorded October 2 and published October 3, 2026. He said the story had been "breaking the internet over the last 24 hours." Then he played a demo. In it, a user plays Simon Says with the avatar, asks for a thumbs-up and has it keep track of a soldering iron's 12-second warm-up. The panel was less interested in the trick itself than in what it proves, and in what should replace the Turing test.
What Tavus built and how it tested it
Diamandis relayed that Tavus calls Griffin the first "human interaction model," and he stressed one feature: it is "fully duplex." That means it listens and talks at the same time, the way people interrupt, nod and react while someone else is speaking. Older avatar systems usually take turns: listen, process, reply. According to Tavus, Griffin keeps processing the caller's audio and video while it responds with speech and generated video.
On the show, Diamandis summarized the headline result as 48% of people mistaking the avatar for a human, compared with 3% for earlier systems. Tavus's own description of the study adds conditions that matter. Participants were told they would have a one-minute call with another study participant about the coming year. They were asked whether their partner might have been artificial only later, after other survey questions, and were then debriefed. Under that setup, 26 of 54 participants believed Griffin-Lite was human, about 48%. Tavus's previous system fooled 1 of 41, or 2.4%. Participants also rated the conversations on seven-point scales: 5.4 for naturalness, 5.6 for trustworthiness and 4.9 for conversational flow.
So the result comes from short calls, with people who had not been told to look for a machine, testing the lighter Griffin-Lite version. Diamandis read out how Tavus pitches the product: "a tutor for every student that notices when they're lost," and an elder-care companion that listens. He also noted that Tavus says Griffin needs safety work before a public release, because it is the first model that can be mistaken for a real person. The company's announcement describes a preview for trusted testers and more safety and disclosure work before customers get it.
Diamandis also said Griffin is number one on NVIDIA's benchmark for full-duplex AI video. NVIDIA's VideoFDB research page is consistent with that: its leaderboard puts Griffin-Lite first among the non-human systems on both the overall perception and overall generation scores. It also shows where Griffin-Lite still falls short of people. The benchmark uses 237 annotated clips of natural two-person calls, and an AI judge scores them. Griffin-Lite scored 3.73 out of 5 on perception, against 4.20 for the human reference, and 3.83 on generation, against 3.92. The biggest gap is timing. On perception, its responses were well timed 73.8% of the time, with a median latency of 2,232 milliseconds. People managed 90% and 1,400 milliseconds. On generation, it scored 62.8% and 1,892 milliseconds, against 78% and 900 milliseconds for people.
"This was expected to happen"
Salim Ismail, founder of Open ExO, was unmoved by how lifelike the avatar looked. "I'm less kind of floored by the imagery and the mimicry of it. This was expected to happen," he said. He wanted to know whether such a system can do something "useful enough that I don't care whether it's human or not."
Diamandis saw a business use: digital co-workers that pop up on Zoom, take calls, texts and Slack messages, and work as what he called "a mechanism for a future virtualized company." Dave Blundin, founder and general partner of Link Ventures, said his reaction matched Ismail's: for anyone who uses AI models every day, the demo is a "so what," since the components already existed and only had to be stitched together. But he expected it to show the mainstream world that the Turing test has been crossed.
Generated faces, China and service jobs
Alexander Wissner-Gross, a computer scientist and founder of Reified, started with a joke. Diamandis "should probably ask me to touch my face," he said, borrowing the demo's Simon Says check. Diamandis played along.
Wissner-Gross then argued that Chinese labs had previewed this direction. He pointed to a streaming model from Alibaba's lab, which in the transcript appears as "one streamer." The likely reference is the Wan team's Wan Streamer v0.1, described on June 24, 2026. He said Chinese labs "remain over-invested relative to the Western labs" in generative video and interactive video models. According to its developers, Wan Streamer uses a single model to take in and produce text, audio and video together. They report video at 25 frames per second and about 200 milliseconds of response delay inside the model, or about 550 milliseconds including network delay, though the version 0.1 demos run at a rough 192p resolution. Wissner-Gross said the streaming model is already out and open source, while Griffin is "not really out yet and definitely not open source."
He called the short term hard to predict and the long term easy. He expects "magic mirrors": fully generated, real-time, interactive video in which whole scenes of people are created on the fly. "It's possible right now, but the latency is high," he said. He doubted the cost-performance case for now, asking how much revenue such models generate per token (a token is the small unit of text or data a model processes and is billed by) and whether they come anywhere close to the frontier set by code generation. "Doubt it," he said. But if generating people for a Zoom meeting or a podcast becomes "so absurdly inexpensive," he asked whether that would still matter.
His answer pointed at jobs. Many human service jobs require a face, a voice and interactivity, he said, and could probably be "completely automated away" by models that until now could only interact through text. He called the technology "transformative for the service sector." He also hoped Tavus and other American developers of video models would take a page from it and start competing with China.
A test no machine can grade
Diamandis put a question to Richard Socher, co-founder and CEO of Recursive and founder and CEO of You.com. In his book The Eureka Machine, Socher argues that AI becomes superhuman where answers can be verified. Does fooling human judges count as a scientific milestone of that kind?
"Not quite," Socher said. When a person has to decide whether the other side is human, he said, the test is hard to scale and "you can't quite automate it completely."
He turned instead to risks. Banks run identity checks called "know your customer," or KYC. Socher said a technology like this calls for going further, to "know your use case." "You don't want this technology to be in Zoom pretending to be the CEO," he said. He recalled "famous stories" in which someone was fooled into wiring about $20 million after a fake Zoom call with five other executives, and said Zoom and Google Hangouts will probably need countermeasures to check whether a participant is a real person. "Scams are going to proliferate on this," Diamandis said.
Diamandis added a personal appeal: if your parents or grandparents are alive, spend hours recording their stories on video and audio. He said that data is what you need to build "a super high resolution lifelike avatar" for future generations.
After the Turing test
Blundin asked what the next milestone should be. He offered the "Demis Hassabis test": give an AI only information from "1910 and prior" and see whether it can rediscover E=mc². He admitted this would be hard to measure, because it is difficult to make sure the system isn't cheating with later knowledge. He asked Socher whether there is a "Richard Socher test."
Socher said he had come up with an "anti-Turing test," because the old test has "completely flipped." To find out whether a human is on the other side, you now ask something no human could do. If you ask for a complex web app and get back 50,000 lines of code ten seconds later, he said, "you kind of know it was not a human."
Wissner-Gross objected: "it could throttle itself," meaning the AI could simply slow down. Socher said that is the point. The Turing test mattered because being as smart as a human was hard, he said. Now a machine just has to "throttle yourself down to human level" to pass, which makes it "useless as an inspiring test for intelligence."
Wissner-Gross offered his own CAPTCHA, the kind of test websites use to tell people from bots, saying "this is not a recommendation." The best way to tell whether you are talking to a text chatbot, he said, might be to ask it a question about chemical, biological, radiological or nuclear weapons and see whether it can respond. His guess: "Probably not."