TwelveLabs Pegasus 1.6 Turns First-Person Video Into Robot Training Data

Teaching a robot to fold laundry starts with someone watching hours of footage of humans folding laundry, then writing down every move. TwelveLabs wants to take that second job off the table.
On 6 October, the video intelligence company released Pegasus 1.6, a model with native support for egocentric video. That means footage shot from the first-person point of view of the person doing the work, whether that is someone cooking, assembling parts on a factory line, or operating a robot remotely.
Why first-person video is a different problem
First-person footage carries different spatial cues and motion patterns from fixed overhead or third-person camera video. It includes motion blur, shifting lighting, and objects that appear and disappear as the wearer moves. Models trained on general video data tend to handle it poorly.
Pegasus 1.6 is built to take that messy input and extract the details that matter for teaching a robot how to move, grasp, or navigate. The model supports five workflows directly.
Action segmentation and labeling generates timestamped labels for tasks, steps, objects, and hand-object interactions from raw video, mapped to a domain-specific taxonomy. Dense caption labeling produces descriptive language for spatial relationships, scene context, and hand-object interactions, aimed at training language-conditioned robot policies. Quality scoring filters out low-quality clips by evaluating action clarity, framing, and stability before footage reaches human reviewers. Search and curation surfaces rare events and long-tail scenarios across a large video repository through natural-language queries. Consent and compliance flagging detects faces, bystanders, and sensitive on-screen or paper data before footage enters downstream pipelines.
The labor math
The number that explains the product is the manual labeling cost. Labeling egocentric footage by hand can take 70 to 155 hours of human time for every hour of video. Run that against a single two-hour clip and a human annotator is looking at weeks of work for one file.
Pegasus 1.6 ships with a 261,120-token context window, large enough to process videos up to two hours long. Access runs through the TwelveLabs API, priced at $1.75 per video hour, $3 per million image-input tokens, and $15 per million output tokens, which is double the rate of Pegasus 1.5.
One caveat for teams planning large jobs: batch analysis is still handled by Pegasus 1.5, so high-volume bulk processing has not fully moved to the new model yet.
The distinction between understanding and doing
The company's CEO, Jae Lee, described the immediate opportunity as training infrastructure. Understanding an action is only part of teaching it. Pegasus can describe limb movements and provide a rough understanding of trajectories, but it does not supply all the information required to execute them. Pressure, touch, and precise control remain separate problems.
That is an honest framing of the boundary, and it is worth holding onto. Robotics teams must combine video-derived information with other data and build the systems that translate it into action. Pegasus sits upstream, feeding the training process rather than replacing the robotics stack.

Lee offered one indication of processing scale, recalling a run that handled roughly 17 years' worth of first-person footage in about 18 hours without a failed video. He did not specify the computing resources or evaluation conditions, which makes that an anecdotal throughput claim rather than a reproducible benchmark. The company's technical materials describe internal evaluations of action identification and recognition of the performing hand, but supply no numerical scores or reproducible comparison against competitors.
Where it sits against the alternatives
The competitive field spans several parts of the robotics development pipeline. NVIDIA's Cosmos Curator overlaps directly on data preparation through filtering, annotation, and duplicate removal, and its open tooling gives teams an alternative to a managed service. Encord's integration of NVIDIA Cosmos Reason 2 and Embed competes for the annotation and curation work, generating preliminary labels for human review with behavior-based search. Google's Gemini Robotics ER 2 overlaps in interpreting video and tracking task progress, though it emphasizes planning and coordination, handing motor execution to lower-level action models. Black Forest Labs and mimic's FLUX-mimic uses a video-model foundation to produce robot actions, addressing execution rather than data preparation.
The billing units and product scopes differ across these tools, so their rates do not settle which is cheapest for a given workload. Pegasus bills video by duration and text output by tokens, which suits teams whose bottleneck is labeling rather than action generation.
What the release does and does not establish
Pegasus 1.6 puts a concrete product behind TwelveLabs' video-understanding research, and it targets a real bottleneck. The path from raw human footage to usable robot training data is genuinely slow and expensive, and automating part of it would matter to anyone building embodied AI.
What remains unsettled is whether the model delivers meaningfully better results than existing alternatives. No third-party benchmarks, pricing comparisons across the full field, or published customer results accompanied the announcement. The clearest thing to watch next is whether robotics developers publish results from using Pegasus 1.6 in actual labeling workflows. That evidence would move the conversation from announcement to assessment, and it is the only kind that settles a claim about labor saved.
Why human video is the base of the pyramid
The strategy behind Pegasus 1.6 rests on an argument about where robot behavioral knowledge comes from, and it is worth stating plainly.
Most of what people know about doing physical work has never been captured in a form a machine can learn from. A cook adjusts a grip without thinking. A worker recovers from a dropped tool by shifting weight and reach in a way nobody writes down. These adjustments are the difference between a robot that performs a motion and a robot that completes a task.
Human video is one of the richest sources of those adjustments, and it is available in enormous volume. The bet is that a broad base of behavioral knowledge extracted from human footage, adapted to specific robots through teleoperated demonstrations, is a faster path to capable machines than starting from scratch for each task.
That is a development strategy rather than a demonstrated guarantee, and the CEO said as much. More video does not automatically produce a broadly capable robot, and translating an understood action into an executed one remains a separate problem involving force, touch, and control.
The consent dimension
One of the five workflows Pegasus 1.6 supports is worth singling out, because it points at a governance question that comes with the territory. Consent and compliance flagging detects faces, bystanders, and sensitive on-screen or paper data before footage enters downstream pipelines.
That feature exists because egocentric video is captured from the perspective of a person at work, and that framing captures more than the task. It catches a badge, a screen, a face in the background, a document on a desk. For a robotics team collecting footage at industrial scale, the flagging step is what keeps a training pipeline from ingesting material it should not.
The inclusion of that workflow alongside the labeling ones is a quiet acknowledgment that the data pipeline has a compliance surface, and that the surface is hard to manage by hand. For teams in regulated industries, the flagging capability may end up mattering as much as the labeling speed, because a pipeline that cannot screen its inputs cannot be used at all.
What scales, and what does not
TwelveLabs frames the release as a move from small, carefully curated datasets toward a scalable pipeline built on real human experience. The honest caveat is that scale in the labeling step does not automatically deliver scale in the training step.
Even a fast, cheap labeling pipeline produces data that still has to be paired with the touch and control signals a robot needs, and still has to be adapted to a specific machine's embodiment. Pegasus 1.6 removes one bottleneck, which is the labeling labor, and does nothing about the others. That is a useful contribution rather than a complete solution, and the teams that benefit most will be the ones whose labeling step was the binding constraint. Whether that describes most robotics teams is exactly the question that published results from actual deployments would answer.
Related articles
C2PA and SynthID: After the EU AI Act, Media Provenance Is an Engineering Task
Provenance tells you where a file came from. It does not tell you whether the file is honest.
Qwen-AgentWorld Puts Seven Environments Inside a Single Language World Model
Agent capability is increasingly limited by the environment around the model, not the model's raw reasoning score.
The Enterprise Agent Plumbing Race: Ampersand and Restate Fund the Read-Write Layer
Neither company builds a model. Both build the plumbing that lets an agent actually do something inside a system a business already runs.
Reactor Raises $74M: Video World Models Need a Runtime of Their Own
A round of this size, backed by the company that also designs the GPUs, tells you where the money thinks the next bottleneck is.