Pegasus 1.6 by TwelveLabs: Transforms egocentric video into robot training data
The model is available through the TwelveLabs API. It supports five functions including action segmentation and labeling, dense caption labeling, quality scoring, search and curation, and consent and compliance flagging. Quality checks can cover obstructed views, unstable cameras and unclear actions. One limitation: batch analysis is still handled by Pegasus 1.5, so bulk jobs have not moved to the new model. Pegasus 1.5 already divided video into timed segments and returned structured descriptions. That release came about six months earlier. Version 1.6 adds training for first-person footage, which the company says earlier versions handled poorly. According to Korean press reports, a U.S.-based robotics data company is already using the model to refine hundreds of thousands of hours of video daily. This is TwelveLabs' first product aimed at physical AI. For robotics founders, the bottleneck has moved from collecting footage to labeling it. TwelveLabs puts manual labeling at $4–$22/hour of footage. If automated labels hold up, the cost of building a manipulation dataset drops sharply, and annotation vendors face pressure on price. Teams should test the output against human-labeled samples before relying on it.