Skild AI Unveils S1 Robot Model to Learn Tasks from a Single Video
- Skild AI has launched its S1 robot foundation model, utilizing in-context learning to help industrial robots learn unfamiliar, long-horizon tasks from a single video demonstration without requiring task-specific...
- Manufacturing floors and warehouses frequently change layouts and introduce new products, forcing traditional robots to undergo extensive reprogramming and data collection.
- The S1 model was developed and trained on NVIDIA AI infrastructure as part of a broader collaboration spanning synthetic data generation, simulation, and physical AI deployment.
Skild AI has launched its S1 robot foundation model, utilizing in-context learning to help industrial robots learn unfamiliar, long-horizon tasks from a single video demonstration without requiring task-specific retraining. According to company disclosures, the model addresses manufacturing bottlenecks by interpreting video prompts to execute tasks such as plant potting and kit assembly.
Task Learning Through Single Video Prompts
Manufacturing floors and warehouses frequently change layouts and introduce new products, forcing traditional robots to undergo extensive reprogramming and data collection. To bypass that cycle, Skild AI designed the S1 model to accept a short video recorded by an operator as a prompt, mapping the demonstrated intent, objects, and sequence directly into physical robot actions. According to company disclosures, the system enables robots to handle unfamiliar, multistep tasks lasting up to 10 minutes—including pour-over coffee brewing and pancake making—without updating model weights. In plant-potting tests cited by Skild AI, engineers moved from recording a video demonstration to autonomous execution on hardware in 11 minutes. When evaluated on novel multistep tasks, robots powered by the S1 model succeeded approximately 66 percent of the time at each step. This compares with a 9 percent success rate for similar AI systems evaluated under comparable conditions, representing a more than sevenfold improvement according to vendor-reported metrics. Furthermore, the company estimates that a single video example provides utility roughly equivalent to 380 hands-on training examples, which can otherwise take human operators between 50 and 100 hours to collect.
Infrastructure and Factory Deployments
The S1 model was developed and trained on NVIDIA AI infrastructure as part of a broader collaboration spanning synthetic data generation, simulation, and physical AI deployment. Deepak Pathak, cofounder and CEO of Skild AI, stated in company announcements that learning by experience rather than preprogramming marks a major shift in robotics. Learning by experience, and not preprogramming, is the step change that has happened in robotics. NVIDIA Isaac Lab and NVIDIA Cosmos technologies help Skild create the scalable, diverse experience its robots need to learn across many scenarios and embodiments. Skild trains its robot brain inside physically based virtual environments using NVIDIA Omniverse libraries and the Isaac Sim framework, applying reinforcement learning in Isaac Lab powered by the Newton physics engine. Ten months after its initial commercial deployment, Skild AI has reached a $100 million annual revenue run rate and established more than 60 deployment partnerships across manufacturing, logistics, inspection, security, and food preparation. Among these initiatives, Skild, NVIDIA, and Foxconn are deploying the Skild Brain on dual-arm manipulators to handle high-precision assembly for NVIDIA Blackwell systems, executing tasks such as installing limit blocks and fastening 16 screws while adapting dynamically to physical disturbances on the production line.

