Figure Helix 2.5 Lets Humanoid Robots Do Chores in 30 Unseen Homes With 56% Success

Helix 2.5 marks progress toward humanoid robots that can work outside carefully prepared environments. But a 56% full-task success rate also shows that reliable home robots remain a work in progress.

Figure Helix 2.5 Lets Humanoid Robots Do Chores

Figure introduced Helix 2.5 on September 17, 2026, saying its humanoids can perform household chores in homes they have never seen before without collecting new data or fine-tuning for those specific homes.

The company tested the system in 30 Bay Area homes. Robots made beds, folded towels and tidied living rooms. Figure reported a 56% full-task success rate, compared with 9% for a policy trained without Index pretraining.

What Figure Tested

Helix 2.5 was pretrained on Index, Figure’s large-scale dataset of human behaviour. The foundation model was then adapted to three chores: tidying living rooms, folding towels and making beds.

The robots were evaluated in 30 homes that were not used for task training. Figure says no evaluation toy, towel or bedding appeared in the task-specification data. The robots also used furniture already present in each home.

Success required completing the entire task. Tidying meant placing all 13 to 15 scattered toys into a basket. Towel folding required folding the towels and putting them in a basket. Bed making required completing the bedding arrangement under Figure’s predefined criteria. Partial completion did not count.

What “Zero-Shot” Means

Figure’s use of “zero-shot” does not mean the robots had never been trained to make beds, fold towels or tidy rooms.

The three behaviours were specified using fine-tuning data collected elsewhere. Zero-shot refers to the evaluation homes and manipulated objects. The robots received no additional adaptation after arriving in those homes.

That makes the test a measure of whether previously learned physical skills can transfer to unfamiliar environments.

How Well Did Helix 2.5 Perform?

Humanoids Daily, citing Figure’s evaluation chart, reported 237 successful trials out of 420, or about 56%.

Bed making recorded 94 successes in 140 trials, or 67%. Towel folding reached 87 of 140, or 62%. Living-room tidying was the hardest task at 56 of 140, or 40%.

Figure compared the Index-pretrained policy with one trained using the same task-specific data, architecture and evaluation setup but without Index pretraining. The company reported 56% success for the pretrained model and 9% for the baseline.

Humanoids Daily noted that Figure’s chart shows 35 baseline successes in 420 trials, equal to about 8.3%. Figure’s written announcement states 9%, so the difference appears to be a small reporting or rounding discrepancy.

Why the Result Matters

The main challenge is not teaching a robot to fold one towel or make one bed. It is getting the same learned behaviour to work across different rooms, furniture and objects without retraining for every location.

Figure says Helix 2.5 used half as much task-specific adaptation data as a representative Helix 02 policy while generalising across 30 unfamiliar homes.

The company also showed examples of robots recovering from mistakes by changing stance, repositioning or moving around furniture before continuing a task.

Still, the reliability gap remains significant. About 44% of the 420 trials did not end in full task completion under Figure’s pass-or-fail criteria. The evaluation was also conducted by Figure rather than an independent testing organisation.

No Consumer Launch Yet

Figure 03 was designed with home use in mind, but Figure has not announced a public consumer price, retail ordering programme or general household delivery date.

That means Helix 2.5 is currently more relevant as a robotics research and development milestone than as a product consumers can buy.

Figure Is Scaling Data and Compute

Figure is investing heavily in the infrastructure behind Helix. On September 3, it announced a partnership with Nscale involving an initial $3.5 billion compute commitment and plans for up to 100,000 GPUs on Nvidia’s Vera Rubin platform.

Figure also said Index is generating about 35 minutes of new human-experience data every second.

The next test will be whether more data and computing power can raise reliability, expand the range of household tasks and produce results that can be independently reproduced.

Sources: Figure, Helix 2.5: Zero-Shot 30-Home Generalization; Humanoids Daily, Figure’s Helix 2.5 Takes on Chores in 30 Unseen Homes; Figure, Figure and Nscale Sign Strategic Partnership

Written by Liam Hisona
Published: September 18, 2026, 8:35 PM PHT

Liam is just starting...

Leave a Reply

Your email address will not be published. Required fields are marked *