Industry Brief “Skilld AI” just released a post on how their model is demonstrating in-context learning, the ability of their model to learn directly from pre-training and with few examples The News Skild AI has announced the launch of S1, its flagship robotic foundation model, which introduces 'in-context learning' capabilities for robots. This new model can learn and execute previously unseen, multi-step tasks up to 10 minutes long from a single video demonstration, eliminating the need for extensive fine-tuning or post-training. What It Is S1 is a robotic foundation model built from the ground up as an in-context learner. Unlike traditional Vision-Language-Action (VLA) models that often rely on language prompts, S1 uses a single egocentric human video as its prompt. It then generates robot actions to complete the task in various environments and embodiments. The model achieves this by composing skills learned during pre-training or creating new ones, enabling it to perform complex actions like making coffee, potting plants, or frying pancakes, even if these specific tasks were absent from its pre-training data. Key Highlights Single-Video Learning: S1 learns new tasks from just one video demonstration, acting as an 'in-context learner' akin to how large language models respond to text prompts. Long-Horizon Task Execution: The model can execute tasks lasting up to 10 minutes and involving dozens of steps, remembering every component and assembly process in detail. Superior…
Skild AI Unveils S1: A Robotics Foundation Model Learning Complex Tasks from Single Video Demonstrations
From the industry · 01/08/2026