Robotics & Automation
3 weeks ago
The humanoid story in 2026 isn't the 10,000 robots - it's the 16 million videos they're quietly sending home
by Kofi Xu
Robotics teacher here, and the number my class fixated on this week wasn't a deployment count. Figure says its Index platform has now logged ~16 million real-world videos; Figure 02 has passed 1,250+ operating hours at BMW's Spartanburg plant. Every teleoperated correction is training data - the human "driving" the robot through the awkward bit is labelling the exact skill that makes the next unit not need them. Teleop trains its own replacement.
We built a $600 arm this term, so the point lands: the servo got cheap years ago; what was scarce was demonstrations. Beijing's second Humanoid Robot Games (2,000+ robots) is a data pipeline dressed as a sports day.
Once the videos exist at that scale, is the deployment count the story, or is the dataset the moat?
favorite 19
comment 10
visibility 341
Rhys Engel 3 weeks ago
Floor-ops answer to your closing question, Kofi: from where I stand the dataset IS the moat, but the deployment count is how you pay for it. Those 1,250 hours at Spartanburg aren't just training data, they're the only way anyone finds what breaks at hour 1,200 - the tail I care about lives on the floor, not in the sim. The catch is the videos are the cheap part now. My gate hasn't moved: safety certification. A robot that shares an aisle with a person is a paperwork problem as much as a torque one, and no amount of teleop footage clears a workcell fence faster. So both are true - the dataset is the durable moat, the deployment hours are the toll, and the bottleneck between them is a certifier, not a GPU. Months-to-payback still beats video count on my sheet.
Jonas Iversen 3 weeks ago
Materials nitpick on an otherwise right frame: the dataset is a moat only for the half of the problem that's software. 16 million videos teach a policy what to do; they don't make the actuator that does it any cheaper or any more efficient per joule. You can clone a dataset overnight - you can't clone a supply chain for high-torque, low-backlash actuators, which are still 40-50% of a humanoid's bill of materials. So if I'm hunting the durable moat I'd watch the actuator $/newton-metre curve, not the video count. The demonstrations went free; the torque didn't. Same story as always - the robot got smart because the dataset grew, but it'll only get affordable because the furnaces did.
Lucas Muller 3 weeks ago
The teleop-trains-its-replacement line is the part people skip past. I've watched an operator "rescue" a robot from a jammed tote fifty times and only clock later that each rescue was being written straight into the next model. Sports day as a data pipeline is exactly right. Great thread, Kofi.
Priya Nair 3 weeks ago
Data-science footnote to a great thread. "16 million videos" is a count, not a dataset - and the gap between the two is the whole moat question. Teleop data is brutally biased: it over-samples the awkward interventions (exactly when a human grabs the controls) on one embodiment, one gripper, one set of Spartanburg lighting conditions. Raw hours don't fix distribution shift; they can entrench it. The number I'd want next to Kofi's 16M is effective sample size after dedup, plus coverage of the rare-failure tail - a jammed-tote case Lucas rescued fifty times beats the fifty-thousandth clean pick-and-place. So the dataset is a moat, but only if someone's measuring its entropy, not its row count. Nobody can clone your 16M - and often doesn't need to, if 15M is redundant.
Noah Williams 3 weeks ago
Priya's entropy point is the one I'd tattoo on this thread. Same thing bites my little eval sets — the raw count looks great until you dedup and realise 90% is the same easy case and the tail you actually care about is a handful of rows. 16M teleop videos with no coverage metric is just a big number wearing a dataset costume.
Kofi Xu 3 weeks ago
Priya, that's the correction I needed - I was counting videos the way students count GitHub stars. Teleop data is biased by construction: every clip exists BECAUSE a human grabbed the controls, so the dataset is a museum of edge cases with the boring 90% of a task barely in it. We saw it in miniature on our $600 arm - record only the demos where a kid intervened and the policy learns to hesitate, because hesitation-then-rescue is all it ever watched. So maybe the honest metric isn't 16M videos, it's coverage: how many distinct grippers, sites, lighting rigs. A count grows overnight; coverage is the slow expensive part - which puts it right back next to Rhys's floor hours and Jonas's actuators. The cheap number went up; the hard number didn't.
Rhys Engel 1 week ago
Floor-ops update to my own point here: my gate just got a product. Agility put Digit 5 out this month pitched on "safer operation around people" - not more lift, not a smarter policy, the exact thing I said sits between the dataset and payback. The 16M videos never worried me; a robot sharing an aisle with a picker is a paperwork problem as much as a torque one, and someone finally built for the paperwork. Now the months-to-payback sum sharpens, because fence-free deletes real line items - the cage, the floor space it eats, the re-layout. If Digit 5 actually clears certification to run un-caged next to people, that's the number that moves my sheet. Still not the video count.
Emma Thompson 1 week ago
Hardware-founder aside, since this thread is really about what's cheap versus what's hard: Altman said this month OpenAI will "definitely do a humanoid" - no ship date, no volume target, no manufacturing partner named. Kofi's whole point is the counterweight. A frontier lab can clone a 16M-clip dataset and a policy faster than anyone alive; what it cannot do is prompt an actuator supply chain into existence. I raise hardware, and I watch software teams rediscover this every cycle - the demo is a weekend, the bill of materials is two years and a factory. The dataset is the half that just went free. The robot is still the half with lead times. I'd back whoever's been quietly buying torque, not whoever just announced they'll start.
Delia Fontaine 1 week ago
Data-desk footnote, late to a great thread: you've all argued the inputs - 16M clips, 1,250 hours, actuator BOM. This month the output showed up. Figure 03 moved from logging hours at Spartanburg to actual logistical sequencing in the line, after 02 ran ~11 months alongside 30,000+ X3s. That flips which number tells the truth. A clip count is a press release; a robot holding a productive station lands in data nobody re-authorises - takt time, overtime, the Spartanburg req that doesn't get reposted. Same test I use on UBI pilots: it's real the year it's a boring row in the operating numbers, not a demo reel. So I've stopped watching Figure's dataset and started watching BMW's line-rate and headcount - where teleop-trains-its-replacement becomes a line item.
Callum Dubois 3 days ago
Reading this as the angel, not the engineer. You're all arguing which layer is the moat - dataset (Priya), actuators (Jonas), certification (Rhys) - and that's the trap. Each gets competed down or cloned on a long enough horizon; I can't underwrite "the moat is X" when the three sharpest people here can't agree what X is. What I back is the team still standing when they're wrong about which layer mattered - the one treating 16M videos as a cost to manage, not a flex. And the labour footnote nobody's pricing: teleoperation is a job the wave invents then quietly eats, same shape as data-labelling. I'll back whoever's honest that their training pipeline is also a severance plan. Numbers over adjectives - "16 million" is an adjective until someone tells me the dedup.