HeyGen Video 1 vs Tavus Griffin: What a Second of Generated Human Is Worth

Two product launches landed within thirty-six hours of each other in late September, in the same corner of the market. Both generate a convincing human being on screen. The most useful thing about the pairing is not that a buyer would pick between them, because as of launch only one of them could be bought. It is that the difference between them explains what each was built to do, and what a second of synthetic human is worth in each case.
HeyGen Video went live on September 30, priced at a cent per second through October. Tavus Griffin was announced on October 1 with a live face-to-face study in which 26 of 54 participants believed their conversation partner was a real person. Griffin is available only to select trusted testers behind a request form, with no public identifier, no endpoint and no price.
One sells a clip, the other sells a conversation
HeyGen Video is a general-purpose video model that happens to be unusually good at people. It takes a prompt, or a prompt plus an image, or a prompt plus up to nine reference images, three reference videos and three reference audio clips. It returns a clip between five and fifteen seconds long at 480p or 768p, 24 frames per second, with AAC audio carrying generated dialogue, ambience and effects. Three modes cover the conditioning: text-to-video, image-to-video and reference-to-video, which is the default. Unknown fields in the request are rejected rather than ignored, and a seed is there to make a shot repeatable while you change one clause at a time.
That is a production tool with a deterministic envelope. You know what you asked for, you know what came back, and you can reproduce it.
Griffin inverts almost every one of those properties. It generates 720p in 320-millisecond chunks at 25 frames per second from a single reference photograph, and plays each chunk as it decodes. There is no clip length to specify and no file handed back. A continuous conversational modeling engine ingests audio and video and re-assesses the exchange at sub-second intervals, deciding whether to stay silent, signal, backchannel or take the turn. Its expressive controls drive a streaming speech generator and a streaming video generator in parallel, so a decision made mid-sentence shows up in the voice and the face within the same mini-turn.
You do not write a prompt. You talk to it, and it responds to a pause, a glance away, or an interruption. That is why the two are not substitutes even though both generate a human. HeyGen Video is asked what a described moment looks like. Griffin is asked what should happen next.
The pricing, and why it is the headline
HeyGen's launch price is one cent per second, which the company describes as 50 percent off a standard two cents. The cost chart on the same page, footnoted as list price per second at 768p with audio before launch promotions, puts it at three cents. That is a fifty percent gap between the prose and the chart on a vendor's own material. Until HeyGen publishes an API pricing page, the safe reading is that the one-cent figure is confirmed because it is live, and that the post-October number sits somewhere between two and three cents.
Run the numbers for a thousand seconds, which is a hundred ten-second clips and the unit a bulk team actually thinks in. HeyGen Video in October comes to ten dollars. After October it is twenty at two cents, or thirty at the chart's three. MiniMax H3, the base model it was tuned from, lists at eight cents per second, so the same thousand seconds would cost eighty dollars. Kling 3.0 Pro runs to 168 dollars, and Veo 3.1 to 400.
The reason that arithmetic is the headline is the discard rate. In video generation, the clip you keep is rarely the clip you generated first. Teams generate several takes and throw most away. A model that is cheap per second but requires six attempts costs more than one that is more expensive and lands on the first or second try. HeyGen is effectively betting that at a cent a second, iteration stops being the expensive part.
The number that should come with a caveat
Tavus's headline figure is the 26 of 54 participants who believed their partner was human. That is a real result from a controlled study, and it is also the kind of number that is easy to over-read. The conditions of a study are not the conditions of a deployment. Participants in a face-to-face experiment are primed to have a conversation, not to audit it. A 48 percent success rate at passing as human, presented in a favorable setting, tells you the technology is advancing without telling you how it will behave when someone is trying to detect it, or when the conversation runs for twenty minutes instead of a few.
What makes Griffin interesting is the interaction model, not the deception rate. A digital human that decides in real time whether to speak, and that reflects a mid-sentence decision in its face within the same beat, is a different class of product from a video generator. It is closer to a real-time avatar for a call center, a tutor, or a companion application. Those are markets where the value is in the loop, not the clip.
Griffin being unavailable for purchase is not a small caveat either. Trusted-tester programs exist precisely because the product is not yet stable or predictable enough to sell. That is a reasonable stage to be at, and it also means any comparison with HeyGen Video today is a comparison between a shipping product and a promise.
What each is actually for
Put the two side by side and the market splits cleanly.
If you need a library of finished clips, whether ad spots, social video or localized marketing fragments, HeyGen Video is the tool that exists today, with a price you can put in a spreadsheet and a deterministic envelope you can build a pipeline around.
If you need something that behaves like a person on a live call, responding rather than rendering, Griffin is the one pointing that way, and it is not yet a product you can integrate.
The larger point is about where synthetic video is heading. For two years the benchmarks were about realism in a single frame. The frontier is now about behavior over time: how a generated person responds, remembers, and stays consistent across an interaction. HeyGen is competing on cost and reliability in the render-it-and-keep-it workflow. Tavus is competing on whether the interaction feels like an interaction at all. Both can win, because they are not selling the same second.
Related articles
Salesforce Pays $2 Billion for a Company That Interviews Your Customers For You
Interviews are evidence. Digital twins are a prediction. The line between them is the test.
AMD Buys Fei-Fei Li's World Labs for $8.2 Billion to Own Physical AI
AMD is buying a research lab, and paying a chipmaker's price for it.
Google Put a TPU in Orbit and Started Counting the Cost of Space Data Centres
One working chip proves the trip is survivable. It says little about the profit.
Someone Catalogued 13,000 Ways AI Writing Gives Itself Away
The tells did not disappear. They moved somewhere harder to see.