AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Why humanoid robots won’t catch up to human workers any time soon

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:A deep dive into the current state of humanoid robotics.

來源Understanding AI作者: Kai Williams

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Today’s Robot Week article is sponsored by 80,000 Hours, a non-profit that helps early-career professionals make the most of their careers. If you’ve been paying any attention to the robotics world in the last couple of years, you’ve probably noticed that humanoid robots are getting better at an impressive pace. In October 2024, Elon Musk had several of Tesla’s Optimus robots serving drinks at an event unveiling Tesla’s new Cybercab. “Optimus is not a canned video. It’s not walled off. The Optimus robots will walk among you,” Musk said at the event. Then in February 2026, the Chinese company Unitree staged a stunning martial arts performance at the Spring Festival Gala in Beijing. A mixed cast of humans and humanoid robots carried out a perfectly synchronized, fluid dance routine. Robots performed spins, jumps, and even backflips. Unitree robots moving in sync at the 2026 Spring Festival Gala. It was a big improvement over the 2025 show, which featured robots walking stiffly across the stage while waving handkerchiefs. Just last week, at the 2026 World Humanoid Robot Games in Beijing, a robot ran 100 meters in 8.86 seconds, crushing Usain Bolt’s human world record of 9.59 seconds. At last year’s competition, the fastest robot took more than 20 seconds to run 100 meters. Demonstrations like these have impressed a lot of casual observers — and created a lot of anxiety about future job losses. If humanoid robots can already serve people drinks, perform elaborate dance routines, and outrun humans, how long will it be before they put millions of people out of work? But if you talk to robotics experts — and I’ve talked to many in recent months — a more nuanced picture emerges. As Physical Intelligence co-founder Karol Hausman put it, people (including himself) “are not very good at judging progress in robotics or judging what is impressive and what isn’t.” Sure, robots can do acrobatic maneuvers that are “very difficult for a human to do,” he said. But then “something as simple as picking up a Coke can turns out to be very, very difficult.” Some of the most impressive demos of humanoid robots involve someone controlling the robot remotely — a process known as teleoperation. It seems pretty clear this was the case with those Optimus robots in 2024, for example. Tesla’s hardware was sufficient to act as a bartender, but its software wasn’t up to the task. So Tesla apparently hired human operators to control the robots remotely. Tesla’s Optimus robot serving drinks to attendees at Tesla’s “We, Robot” event in October 2024. (Screenshot from Tesla’s official livestream) And while those Unitree robots’ dance moves were not teleoperated, they don’t tell us all that much about the robots’ capacity to do useful work. Most physical labor involves manipulating objects in the real world — packing boxes, hammering nails, flipping hamburgers, and so forth. As we’ll see, training a robot on physical manipulation tasks like these is much harder than training a robot to dance. There are also broader challenges that transcend individual tasks. For example, human workers are extremely flexible — they can perform a wide variety of tasks, and they can learn easily while on the job. So far, nobody has figured out how to give AI robotics models the same capacity for generalization. Today’s most impressive robotics demos involve tasks that take humans several minutes at most. But human workers also perform tasks that take hours — things like “rebuild this car’s engine” or “assemble those kitchen cabinets.” Training a robot to complete longer projects requires building skills unnecessary in short tasks, like the ability to keep track of what’s already been done. Then there are a lot of practical economic and safety concerns that will become obvious once we try to deploy robots in the real world. Robots will need to work for hours without breaking down. They can’t be too expensive to manufacture, train, or repair. They need to be extremely safe to operate in proximity to human beings. It will take many years — maybe even decades — to overcome all of these challenges. So yes, humanoid robots have made a lot of progress in the last few years. But there’s still a long road ahead. Manipulating objects is hard A key challenge in robotics is predicting how the outside world will react to a potential robot action. In this respect, dancing is simpler than most other tasks because (as Bracket Bot CEO Brian Machado told me) “the floor doesn’t do anything.” But while acrobatic robots are impressive to watch, it’s not actually that useful for a robot to dance or do backflips. Most useful work involves interacting with objects that move and change in response to a robot’s actions. “The really, really core unsolved problem in robotics that unlocks 90% plus of the value is manipulation,” Theophile Gervet, president of the robotics startup Genesis AI, told me. Picking up an object doesn’t just change its location, it can also change its shape. And different objects respond in different ways that are hard to model in a general way. Think about the different ways that a pillow, a bag of chips, and a glass of water behave when they are picked up. In September 2025, the roboticist Benjie Holson (formerly Google X, currently OpenAI) announced the Humanoid Olympics, a list of 15 manipulation tasks that he believed would require researchers to “push the state of the art” for a robot to be able to solve. Most of them would be trivial for an eight-year-old child to perform. Three of the tasks involved opening doors. Another was to make a peanut butter sandwich given bread and a closed jar of peanut butter. Perhaps the hardest task on the list for a human to perform would be to peel an orange. To demonstrate the tasks, Holson dressed up in a silver robot suit and took videos. This is a screenshot of a video of him demonstrating the gold-medal door task. Even with tasks this easy for humans, it was an impressive accomplishment when — three and a half months later — the startup Physical Intelligence announced that it had successfully demonstrated 10 of the tasks. Having a robot company “do basically almost all of them in the first three months is wild,” Holson told Scientific American. But Physical Intelligence’s performance came with caveats. The researchers taught the robot how to do these tasks by puppeting a robot over and over until they could fine-tune a model to complete the task. To turn a sock inside out, they trained on 176 successful examples, or around eight hours of data. They peeled so many oranges that the researchers told Holson that the “corner grocery probably noticed the increase in orange sales and the one guy at the company who really liked mandarins was getting pretty tired of them.” The robot took four to 10 times longer than a human to complete almost all of these tasks — while only succeeding 52% of the time! None of this is meant to dismiss Physical Intelligence: its result was a genuine accomplishment. But even on these fairly simple tasks, robots are still far from human-level performance. And it’s still easy to find tasks that are straightforward for humans but entirely beyond the abilities of robots. In January, Holson released a new set of manipulation challenges. While these tasks are more difficult, they are still straightforward for most adults: make a bed, hammer a nail, catch an egg without breaking it. One category of manipulation task in Holson’s new list is worth noting: those that take a long time. Many humans are able to complete physical tasks which take hours — such as putting a bed together or painting a room. But like current LLMs, robots today struggle to complete longer, many-part tasks. Holson included two tasks he dubbed “long horizon”: taking out the trash from a home and making an egg sunny-side up. Neither task took longer than five minutes. I imagine it will take a lot of work to extend the capabilities of robotics models past these several-minute tasks. Current robotics models only have a limited capacity to remember what actions they’ve previously taken — for instance, Physical Intelligence’s most recent model can remember up to 15 minutes using a method to compress its previous observations into text.1 Models also aren’t yet reliable enough to string tens or hundreds of diverse subtasks together. When I asked Gervet about long-horizon tasks, he said he wasn’t worried about it. “I think the job of a robot foundation model company, whether it’s full-stack or not, is to build more low-level five- to 10-minute horizon tasks” rather than to completely solve robotic reasoning. He expects that general-purpose AI companies will solve long-time planning and execution. But it’s not obvious that it will be possible to cleanly separate short-term tasks from longer-term planning. Often, as humans work on individual subtasks, we learn things that cause us to change our overall plan. A system where different models are responsible for long-term planning and short-term task completion may lack our capacity for real-time adaptation. Robot Week special: get 25% off an annual subscription. Generalization Holson announced his second batch of challenges more than seven months ago. As far as I can tell, no one has announced a solution to any of them. Maybe that will change in the coming months. But even if manipulation capabilities increase dramatically, there is an important interlocking challenge: generalization. Robot companies have figured out how to train a robot on a specific task by pouring tremendous effort into that one task. But most individual tasks where this type of high-effort approach makes economic sense have already been solved using conventional automation techniques. So a key bottleneck is whether robot capabilities can generalize to new environments without significant training data. For instance, when I visited China in May, I saw a Galbot robot working in a pharmaceutical warehouse. Its task was to take boxes off the shelf one at a time and place them in a chute for delivery workers. Despite the simplicity of the task, it took the robot around 40 seconds to put each item into the chute. (Galbot said that newer deployments are twice as fast.) Every time a new item was added to the warehouse, Galbot had to train the robot to be able to pick up that specific new item as well — though the company said it only takes five minutes of training. A Galbot semi-humanoid robot grabs an item from a pharmaceutical warehouse in Beijing, China. (Photo by Kai Williams) Basically every humanoid deployment today takes a similar approach. Companies choose a task with enough variability that traditional automation methods won’t work. But there can’t be too much variability, or else contemporary AI methods won’t work either.2 And each deployment typically requires a ton of setup and training effort. Some startups expect that AI will make it easier to deploy robots across a broad variety of tasks. Jagdeep Singh, the CEO of Rhoda AI, told me that the company has “literally over 100 different use cases” in manufacturing and warehousing they’re working on bringing to deployment. When I went to Nvidia GTC in March, several roboticists told me that the most impressive demo they saw was from a company called Generalist. Screenshot from Generalist’s blog post about its GTC demo, running the demo task in the company’s office. Generalist did not do any specific training to transfer the demo task to the GTC exhibit hall environment. Generalist showed a robot inserting and removing a phone from its box using two robotic arms. While the task itself is pretty difficult for current robots, what impressed people the most was how little time and effort Generalist put into the demo. The company claimed that it only had a handful of days to prepare the demo and, notably, did not train the robot to do the task in the exhibit hall itse [truncated for AI cost control]