Skild AI has a robot foundation model that can pick up tasks it has never seen before after watching one video, and the company built it on NVIDIA's AI infrastructure. The model, called S1, launched the week before the September 10 announcement. No price or general availability date has been disclosed.
The setup is simple on the operator side. A person records a video of the task they want done, and that recording becomes the prompt. S1 works out the intent, the objects involved and the order of steps, then maps them onto robot actions. Nothing is retrained and no task-specific post-training is needed, so the robot can attempt jobs that fall outside its pretraining data. The technique is known as in-context learning.
Read more: Apple reclaims world's most valuable company title as Nvidia loses $173 billion in one day
In its own tests, Skild reported that S1 got through about 66% of steps on new multistep tasks, while a comparable AI system managed 9%. A single short video is treated as roughly equivalent to 380 hands-on training examples, which Skild says could take 50 to 100 hours to collect by hand. In one plant-potting test, the gap between the demonstration recording and autonomous execution was 11 minutes. The model handles unfamiliar tasks that run as long as 10 minutes and span dozens of manipulation steps, adjusting when objects shift and recovering from mistakes.
Skild points to jobs such as potting plants, making pancakes, brewing pour-over coffee and assembling kits. In one workflow a robot drives 16 screws while coping with disturbances, a job NVIDIA describes as requiring accurate movement, contact-aware control, step tracking and error recovery. The company says operators can demonstrate new tasks without building a new dataset or running a training job. Where customer agreements allow, what the robots learn in the field feeds back into the broader model.
The collaboration covers synthetic data, training, simulation and real-world deployment. NVIDIA's Cosmos open world foundation models broaden the training data, and Cosmos Curator annotates, filters and organizes it at scale. Omniverse libraries and Isaac Sim supply virtual environments, with reinforcement learning handled in Isaac Lab on the Newton physics engine. The two companies are jointly building GPU-accelerated simulation solvers for touch, grip and the manipulation of solid objects, which the announcement says will soon be released to every developer through Newton. Nsight tools track down training bottlenecks, and TensorRT speeds inference so robots respond quickly in the physical world.
Skild has assembled more than 60 deployment partnerships, and the work spans manufacturing, logistics, inspection, security and food preparation. The company says it hit a 100 million dollar annual revenue run rate 10 months after its first commercial deployment.
The announcement lands in a year NVIDIA's chief executive had already framed as a turning point for robots. Jensen Huang predicted 2026 would be physical AI's ChatGPT moment, and the company used CES to release its Isaac and GR00T foundation models. NVIDIA has also struck a partnership with LG Group covering humanoid robots and next-generation data centers. Skild's CEO, Deepak Pathak, argues the shift is about how robots acquire skills at all: "Learning by experience, and not preprogramming, is the step change that has happened in robotics".
What the announcement does not include is any way for outsiders to buy or try S1. There is no price, no release window and no stated hardware requirement beyond the NVIDIA stack the model was trained on.











