Figure’s Helix 2.5 announcement is a better robotics signal than another polished manipulation video because it publishes a harder test shape. Figure says the humanoid system was evaluated across 30 Bay Area homes it had not seen before, with no evaluation-home data collection, no fine-tuning in those homes, and no adaptation to the objects used in the trials.
The important metric is not whether a robot folded one towel on camera. Figure reports full-task success, meaning every toy tidied, every towel folded, or the entire bed made, with no partial credit. Its Index-pretrained model reached 56 percent zero-shot success, compared with 9 percent for a same-architecture model trained from random initialization on the same task data.
Grey Haven’s read: this is the right direction for judging physical AI. The operator question is not whether a humanoid can perform a scripted subtask. It is how much site-specific data, retraining, supervision, recovery, and maintenance are needed before the system pays back in a real environment.
For manufacturers, care facilities, warehouses, and field-service operators, the lesson is to ask vendors for deployment-shaped metrics. How many unseen sites? How much new data per site? What counts as success? How often does a human reset the task? What happens when lighting, clutter, object variation, or workspace geometry changes?
The next thing to watch is whether these results survive outside carefully selected homes and into commercial workflows with time pressure, safety constraints, and ugly edge cases. If full-task success keeps rising while site adaptation falls, humanoids become a labor-planning question. If not, they remain expensive demo equipment with better marketing.
Source: Figure AI, “Helix 2.5: Zero-Shot 30-Home Generalization,” September 17, 2026.