Robot foundation models
Reading time 6 min readSkild AI

Skild S1 learns new multistep robot tasks from one video without task-specific retraining

Skild AI says S1 can use one video demonstration as context for unseen tasks lasting up to 10 minutes without updating model weights.

By TechniaHQRobot

Skild AI robot performing a multistep manipulation task

Introduction

Skild AI introduced S1 in August 2026 as an in-context robotics foundation model that takes a video demonstration as the task prompt and attempts unseen tasks without task-specific weight updates.

Reported examples include plant potting, pancake making, pour-over coffee and kit assembly. Skild says some tasks last up to 10 minutes and span dozens of manipulation steps; these are company-reported results and should be evaluated with complete-task success and intervention data.

Information verified from official sources available as of September 27, 2026.

In-context robot learning cuts the setup loop

A conventional deployment may collect many demonstrations then retrain a policy. S1 uses the demonstration as context at run time.

Skild reports that one plant-potting test moved from recording the example to autonomous hardware execution in 11 minutes.

Long tasks expose compounding failure

NVIDIA reports about 66 percent success per step in Skild tests on new multistep tasks. A long sequence can fail when any one step fails.

Sequence-level completion rate, recovery count and intervention time would give a fuller view of performance over ten-minute tasks.

Video demonstrations can hide contact information

A single camera view may miss force, depth or hand contact. The model has to infer those details or recover using robot sensors during execution.

Tests with occlusion, moved objects and different viewpoints can show how far one-video learning reaches before another example is needed.

S1 is now also a base for post-training research

On September 23, Skild published a separate physical self-play result that uses S1 as a strong base model for reinforcement-learning post-training in simulation. That work should not be merged with the single-video in-context learning claim because it adds a different training stage.

The distinction is useful for readers: one-video prompting tests whether S1 can infer a task from context without task-specific weight updates, while self-play tests how the base model can improve through additional reinforcement learning.

Limitations and missing information

  • The performance figures are reported by Skild through NVIDIA's article.
  • Per-step success does not equal full-sequence completion.

Conclusion

The strongest follow-up metric is complete-task success across many unseen multistep tasks with recovery and intervention counts included.

Sources and methodology

Share this article

Share the current TechniaHQRobot article page.

Continue reading

Open the latest robotics reporting, Physical AI analysis and hardware notes.

Browse robotics news
Article by @techniahqrobot