The Data Bet: How Robot AI Companies Are Choosing What to Learn From From a million hours of egocentric human video to handheld grippers and zero-shot home deployments — the biggest open question in robot foundation models isn't the architecture. It's the data. Every robotics AI company building a foundation model eventually faces the same question: what do you train on? The answer shapes everything — the hardware you build, the model architecture you choose, the tasks you can tackle, and how fast you can scale. Right now, a handful of well-funded startups are making meaningfully different bets. Some are going all-in on raw human video at internet scale. Others are investing in bespoke wearable or handheld capture hardware to get richer, denser signals. Others are deliberately mixing every source they can get their hands on. And a few are chasing the holy grail: a single model that generalizes to any robot body from any data source. Here's where the field stands. --- The Core Problem: Why Robot Data Is So Hard to Scale Traditional robot learning relies on teleoperation data — human operators manually driving robots through tasks, generating paired (observation, action) trajectories. It works. But it's brutally expensive and, critically, it can't scale to the volume that produces truly general models. Skild AI frames the dilemma precisely: every major source of robotics data makes a trade-off across three axes that matter — hardware proximity (how closely data resembles the…
How are Companies Training their Robot Foundation Models: Human Egocentric Data vs Hand-held Devices
Reports · 18/09/2026