A humanoid robot walked into 30 strangers' homes in September 2026 and successfully made beds, folded towels, and tidied toys more than half the time, without being trained on a single one of those homes beforehand. That is the headline result from Figure AI's Helix 2.5 trial, the most detailed published data on in-home humanoid performance to date.

The result matters for anyone tracking the humanoid robot market, because it moves the conversation from controlled lab benchmarks to real domestic environments. Whether you are watching for investment signals, planning a business deployment, or simply assessing how far consumer humanoids are from being viable used purchases, Helix 2.5 represents a meaningful data point.

What Helix 2.5 Actually Does

Helix 2.5 is a vision-language-action model, a class of AI architecture that fuses a robot's camera feed, natural language understanding, and physical motor commands into a single unified system. Rather than running separate vision, planning, and actuation modules in sequence, a VLA model processes all three simultaneously.

On the Figure 03 hardware, Helix 2.5 receives a camera feed of the room, interprets what it sees, decides what action to take next, and outputs motor commands, all in a tight loop. The result is a robot that can, in principle, read a room it has never seen and begin working through a task without being explicitly programmed for that layout.

The pre-training dataset Figure calls "Index" is central to this performance. Compiled from thousands of hours of human behaviour data across dozens of homes, Index gave Helix 2.5 a broad prior understanding of how household objects look, where they tend to be, and how they are typically manipulated. Without Index, the baseline success rate on the home trials dropped from 56% to 9%, a figure that makes clear how important the training corpus is relative to the model architecture alone.

The Trial Numbers in Detail

Figure tested the Figure 03 robot in 30 Bay Area homes that neither the robot nor the research team had visited before. No environment-specific training was performed. Three tasks were evaluated:

Making beds: 94 successful completions out of 140 attempts, giving a 67% success rate. Bed-making is the most structured of the three tasks because it involves a defined surface and a repeatable sequence of movements. The higher score here reflects that predictability.

Folding towels: 87 successful completions out of 140 attempts, giving a 62% success rate. Towel folding requires the robot to grasp a flat, deformable object and manipulate it through multiple folds without a fixed anchor point. It is a harder manipulation problem than it appears.

Tidying toys: Figure did not publish a per-task breakdown for toy tidying with the same precision, but the overall 56% aggregate across all three tasks indicates this was the weakest of the three. Toy tidying involves the highest variability: objects differ in shape, weight, and position across every home.

The key phrase in every one of these results is zero-shot generalisation. The robot had no map of the home, no knowledge of where the bedroom was, and no dataset labelled for those specific towels or bed types. The 56% figure is therefore not the ceiling of what Helix 2.5 can do in a familiar environment. It is the floor performance in completely unfamiliar ones.

Why This Matters for the Used Robot Market

The used humanoid market is currently negligible in the UK. No humanoid has yet depreciated into a price range accessible to consumers, and failure rates in real environments have been too high to justify business deployments outside controlled warehouse settings. Helix 2.5's trial results begin to change that calculus, though not immediately.

For context, the used robot vacuums available today on the UK secondary market succeed at their primary task at rates above 90%. A Roborock S8 Pro Ultra bought second-hand navigates a real home reliably because the underlying technology has been refined through millions of deployments. Humanoids are perhaps a decade behind that level of reliability for household tasks.

What the Helix 2.5 trial establishes is that the trajectory is real. A 56% zero-shot success rate in 2026, with the baseline starting at 9% before large-scale training data was applied, suggests the limiting factor is data volume and model scale rather than a fundamental hardware barrier. As training datasets grow and model iterations accumulate, success rates will rise.

For buyers interested in used humanoids today, the practical picture remains the same as in early 2026: the Figure 03 and comparable systems are enterprise products with enterprise price tags, and the secondary market for them is essentially non-existent in the UK. That changes when enterprise customers begin rotating stock, most likely from 2028 onwards based on current deployment timelines.

The Broader Context: Where Other Humanoids Stand

Figure Helix 2.5 is the most data-rich published result in the humanoid space, but it is not the only development worth tracking.

Agility Robotics unveiled Digit 5 on 15 September 2026, with the headline capability of operating without physical safety cages alongside human workers. Digit 5 lifts 50 lbs, runs for 20 or more hours per day, and has $300 million in multi-year customer orders. Its deployment target is manufacturing and logistics rather than household tasks, but the cooperative safety architecture is a meaningful hardware advance.

Tesla Optimus continues to operate exclusively inside Tesla factories. No external delivery date has been confirmed, and no independent task accuracy data has been published. The contrast with Figure's transparent home trial methodology is notable.

XPENG IRON began rolling off a production line in Guangzhou on 8 September 2026. At an international price of approximately $120,000, it sits in the enterprise tier and is targeting retail store guide roles in Q1 2027 before broader deployment. The IRON has 76 degrees of freedom, giving it more dexterous hands than most competitors, which will matter for manipulation tasks.

What to Watch Next

The Helix 2.5 trial is a single data release, not a product launch. Figure AI is selling its robots to enterprise customers including BMW, not to individuals. The research value is in what the numbers tell us about the pace of development.

Three signals are worth watching over the next 12 months. First, whether Figure or a competitor publishes home-trial data showing consistent performance above 75% on a standardised task set. Second, whether Agility Robotics or Figure confirm any deployments in UK-based operations, which would indicate European commercial readiness. Third, whether Unitree or any Chinese manufacturer releases a consumer-grade humanoid with verified UK pricing, since Unitree's Go2 Pro quadruped has already demonstrated that Chinese robotics companies can reach the UK secondary market quickly once they choose to.

For now, the Helix 2.5 results are the best evidence yet that in-home humanoid robots are a 2030s product rather than a 2040s one. That window is relevant if you are planning business automation, considering research investment, or simply tracking when the used humanoid market will become worth monitoring in the UK.

For a practical look at which robot categories are already worth buying second-hand today, see our used robot buying guide and the humanoid robot buying guide for used models. Current humanoid profiles including the Figure 02 predecessor are listed in the Robot Database. The question of whether humanoids can genuinely do housework is explored in depth in our analysis of humanoid robots and household tasks.