25 September 2026
Heard In AI

Figure tested its humanoid robot on chores in 30 unfamiliar homes. Vlad Tenev isn't sure he wants one in his house

On September 17, Figure reported that it tested its Helix 2.5 system on tidying rooms, making beds and folding towels in 30 Bay Area homes the robot had never been in before. In Figure's controlled comparison, a version pretrained on videos of people doing everyday tasks completed 56% of full tasks, compared with 9% for a version trained from scratch. The robot had still been trained on each chore elsewhere, and any safety intervention counted as a failure. On the Moonshots podcast, investor Dave Blundin argued that more data could carry skills across very different tasks. Robinhood CEO Vlad Tenev said the machines look too intimidating to have around his children.

A briefing reports one development at a point in time. We may correct or clarify it later; a new development gets a new briefing. How our formats work

Based on Moonshots with Peter Diamandis, episode published 19 September 2026

Making a bed is easy for a person. For a robot in a stranger's house, it is hard: the robot has never seen the bed, the pillows sit wherever someone left them, and the comforter has to end up straight. On September 17, the robotics company Figure said its humanoid robot had started to do jobs like this in homes it had never visited. The next day, on the Moonshots with Peter Diamandis podcast, two of the hosts praised the work. Their guest was not sure he wanted such a robot anywhere near his children.

The episode was recorded on September 18 and published on September 19. Host Peter Diamandis, founder of XPRIZE and Singularity University, played a short Figure video. He was joined by Dave Blundin, founder of Link Ventures; Alexander Wissner-Gross, a computer scientist and founder of Reified; and Robinhood co-founder and CEO Vlad Tenev.

What Figure reported

Robotics researchers call the ability to handle new situations "generalization". The Figure video calls it "the holy grail for robotics". Until now, the video says, Figure had run its Helix AI system autonomously, but it had collected data in each environment where the robot worked. With Helix 2.5, the video says, the Figure 3 robot can work in rooms it has never seen and handle unfamiliar household objects wherever they have been left.

The video shows three tasks: tidying a living room with toys scattered around, making a bed and folding towels. It calls towel folding "really difficult for robotics" because it needs precise handling.

Figure's written report explains the test. The three behaviors were evaluated in 30 Bay Area homes the robot had not been in before. Figure calls the result "zero-shot", which here means the robot received no training in those homes or on the objects it handled there. The robot was not learning each chore from scratch, however. Each task had been fine-tuned (given extra, task-specific training) elsewhere, and one fixed version of the software was used for each task in every home.

Figure used strict scoring. A trial counted as a success only if the robot finished the whole task, not part of it. If a person had to step in for safety, the trial counted as a failure. To pass bed making, the robot had to place both pillows and straighten both corners of the comforter within a time limit.

Learning from people's videos

The main technical claim concerns where the robot's skills come from. Figure runs a program called Index, which pays people to upload videos of themselves doing everyday physical tasks. The robot's software first learns from this human footage, a stage called pretraining, and only then gets practice on specific chores.

To test whether pretraining helped, Figure trained two versions that differed in one way. It kept the design, the task training, the training method and the scoring the same. One version started from scratch and the other was pretrained on Index videos. According to the report, the version that started from scratch completed 9% of trials and the Index-pretrained version completed 56%.

Figure also looked for a pattern familiar from chatbots. A scaling law describes how a model's performance improves predictably as it gets more data and computing power. As the video puts it, language models got steadily better at predicting the next word as developers doubled both. The video says Figure found something similar when its models predicted a robot's next action and the company repeatedly doubled the Index data. It calls these "direct human-to-robot transfer scaling laws".

The report defines this narrowly. Figure trained four models of the same size, increasing the pretraining data eightfold across them. The error in predicting actions fell in a predictable way. For the largest run, Figure says its forecast missed by 0.54% of the measured variation in that error. The video puts it more strongly, saying Figure could predict the result "down to four decimal points" before training began. Either way, the smooth curve describes prediction error during training. It is not a law of how often a robot will finish a chore in someone's home. The video itself says robot learning is not "fully solved yet" and calls Helix 2.5 "the first signs" of more general physical intelligence.

Figure has given different numbers for the program's size at different times. Its August 25 Index announcement described an effort developed over four months. It reported 264,000 app downloads in 108 countries, more than 44,000 weekly active contributors, over 16 million uploaded videos and $15 million paid to contributors. Examples ranged from laundry and cooking to vehicle maintenance. The later video shown on the podcast says Index launched "a year ago" and that more than 90,000 people now contribute every week.

Why Blundin expects skills to carry over

Blundin argued that data is the real advantage. Physical-world data of this kind has never been captured before, he said, and with it a company can "rebuild the model" basically every night. He expects robotics to follow the pattern of language models, where what a model learns from one kind of example helps with others more than people would expect. Blundin said that moving from folding a towel to putting a spare tire on a car sounds like a big leap, but argued that there is "a lot of commonality in physical movement". If Figure extends its lead in data, he predicted a "crazy explosion of capability".

He also praised the demonstration for what it was not. Many robot videos that capture attention on X, he said, are contrived, pre-programmed, pre-scripted or in many cases controlled by a person behind the scenes. He said Brett Adcock, who leads Figure, has been "unbelievably honest" from the outset that there is no human behind his robots. Diamandis recalled an interview the show recorded at Figure's headquarters in January. He said Adcock predicted then that by the end of 2026, humanoid robots would perform unsupervised multi-day tasks in homes they had never seen. Diamandis called the new results "at least the first steps" toward that goal.

Tenev: "like the Terminator"

When Diamandis asked Tenev whether he would get a robot for his home, Tenev hesitated. "I think I wouldn't have one of those in my house," he said. He asked why these robots all "look like the Terminator". He did not want a "scary looking thing" making beds and picking up toys in his children's room, or to run into one on the way to the fridge for a midnight snack. To Tenev, the product had not been figured out yet; what he saw was "a technology demonstration with charts of scaling laws". He could picture such a machine on a construction site building his house. He asked why no one had built a friendly household butler like C-3PO, and doubted many people would buy one to rock their baby to sleep.

Blundin supported the point. He said a lab at the MIT Media Lab studies how children interact with robots, and the children always want them cuddly, furry and friendly, "kind of look like Elmo". Silicon Valley's robot companies, he said, are going the opposite way. Tenev's explanation was that the industry copied Elon Musk. He said Musk had an industrial robot for heavy work in mind, and no one seemed to have asked whether a robot that empties the dishwasher should look very different.

The panel named some exceptions. Wissner-Gross pointed to Apple's work, which he said had been well publicized, on something like a Pixar lamp: a HomePod with a screen on an adjustable robotic arm that can look around. Tenev said that "sounds awesome". Diamandis recommended Sunday Robotics, whose home robot he described as friendly-looking and "quite cheery". He agreed with Tenev that none of the labs had gone in that direction yet.

Sunday's website describes its Memo robot very differently from a walking humanoid. Memo has a stable wheeled base, a low center of gravity and a rounded silicone covering, and owners can choose its colors and hats. It is designed to stay stable without constantly balancing on two legs. Its skills include handling dishes, making espresso and folding socks. Like Figure, Sunday trains on human demonstrations: people wear its Skill Capture Glove while doing chores in their own homes, so no one has to operate the robot remotely inside each customer's home. Sunday says reliability and generalization are still being improved. It is advertising a household beta for late 2026, with general sales to follow after testing.

The customer who wants twelve

Tenev admitted he "could be completely wrong". His middle child, he said, had asked for 12 humanoid robots for Christmas. Tenev asked whether he was worried the robots would take over the house and kill them all. Blundin joked that 12 made a soccer team with a spare. Tenev's practical question to his son was where they would keep them all.

Share this article

Go to the original

Sources & further reading

  1. 01
  2. 02
  3. 03

Connected ideas and articles

From the conversation

Podcast episodes

Moonshots with Peter Diamandis

Robinhood's Vlad Tenev on Tokenizing Everything, OpenAI's 6 Misalignment Reports, Figure's Robot Makes Beds | EP #292

Episode published This article draws on 1:30:27–1:34:42 and 1:36:34–1:39:50 (approximate times)

Article history

Updates to this article

Tags

In a small pretraining experiment, better data beat better architecture — but the Moonshots panel split on how long data advantages last

A Moonshots panel unpacks Dwarkesh Patel and Jerry Han's experiment, which found that improvements in training data delivered a 12-fold compute-efficiency gain between 2019 and 2025 against 3.7-fold for architectures and training recipes — at small scale, on easy benchmarks. The panel then splits over whether a company's proprietary data is a durable advantage, with a $32 billion data-subsidiary valuation on one side and the fate of BloombergGPT on the other.

· Updated 8 min read

Free tokens or owned robots: two ways the panel would share AI's gains

On the Moonshots panel, Alex argued that China's AI loyalty perks are the start of "universal basic tokens": redistributed access to machine intelligence. Just before he spoke, another panelist challenged the idea that AI will become "too cheap to meter." Costs will fall, he said, but people will want far more agents, so budgets will stay real. He also told listeners to grab free tokens when they are offered. Emad Mostaque, who called for tokens for everybody, also proposed 100 million publicly underwritten robots owned by the people. Alex said that sounded like communism and said he would rather see every American own robots individually. Mostaque replied that the government did not need to own the robots, but that ownership had to reach the people somehow.

· Updated 6 min read

Beyond the copilot: career advice from a panel that disagrees about jobs

On Moonshots, Dave Blundin, founder and general partner of Link Ventures, describes a hire in his early twenties who runs his agents entirely by voice, and argues the goal is to manage swarms rather than lean on a single copilot. Host Peter Diamandis's optimistic jobs roundup runs straight into Emad Mostaque's warning that today's hiring is "the turkey before Thanksgiving" and the view of computer scientist Alexander Wissner-Gross that every profession, trades included, is only a question of sequencing.

· Updated 7 min read