Skild AI introduced S1, the company's new foundation model for robotics that can learn previously unseen tasks from a single video demonstration without fine-tuning.
In its Tuesday post on X, Skild AI said S1 uses in-context learning to perform long-horizon tasks lasting up to 10 minutes. S1 can execute these actions in various environments and on different robot embodiments, the company said.
Demonstrated skills include making coffee, potting a plant, and frying pancakes, which were not in the model's pre-training data.
"The first time S1 flipped a pancake, we assumed pancake flipping must have been in its pre-training data," Skild AI said. "We searched our whole pre-training data and found no examples of flipping. S1 inferred the out-of-distribution task from one video prompt."
According to the company, S1 is a "step-change" improvement compared to conventional vision-language-action (VLA) models - it can match the performance of language-prompted VLAs for known tasks, and outperforms existing language-prompted VLAs in learning completely new tasks.
Skild AI estimated that existing VLA models would require 50 to 100 hours of additional data collection and fine-tuning to reach comparable one-example accuracy.
The team also noted that their new model does not blindly mimic the video demonstration, but displays "common-sense" understanding beyond the given prompt. It can withstand perturbations, improvise, or even display more precision than demonstrated in video.
S1 expands on the mobility in-context learning capabilities Skild AI unveiled in LocoFormer last year, which enabled a robot to adapt to disturbances like disabled legs and locked wheels.
Skild AI said it has begun deploying S1 to limited industrial partners and is planning a broader rollout in the coming months.
