Back to feed
News Story
量子位
2 sources

Robot GPT-3 Moment? Generalist AI Releases GEN-1.5, Learning New Actions from 3-Second Demos

Generalist AI has released GEN-1.5, a robot foundation model that enables one-shot learning from short demonstrations, allowing robots to acquire new skills by watching 3-12 second videos without gradient updates or fine-tuning. The model can also combine different demonstrations and transfer skills from simulation to reality, marking a potential 'GPT-3 moment' for robotics.

SynthePulse Insight · AI deep reading

GEN-1.5: One Demo Teaches Robots New Tasks, but Success Rates and Verification Remain Questionable

Version 1 · 1 source

Robotics startup Generalist AI releases GEN-1.5, claiming it can teach robots new tasks from a single demonstration without training. But an average success rate of 59%, a fine-tuned performance of 83%, and a lack of independent verification leave this 'general' capability in question.

  • GEN-1.5 loads 3-12 second demonstrations as 'physical prompts' into the context window, enabling robots to perform tasks without training.
  • The company reports an average success rate of 59% across ten tests, improving to 83% after ten training steps and five minutes of data.
  • The model can chain two prompts, use simulation demonstrations, and partially mimic human hand movements, abilities claimed to emerge spontaneously during eight months of pretraining.
  • Generalist claims to be the first team to achieve in-context learning across a wide range of tasks, but all results come from the company itself and lack independent verification.
Open section navigationCore Mechanism: Physical Prompts and In-Context Learning

Core Mechanism: Physical Prompts and In-Context Learning

GEN-1.5's core innovation is loading demonstration videos as 'physical prompts' into the model's context window, functioning as short-term memory. Users simply provide a 3-12 second demonstration, and the robot can directly perform the task without additional training. This mechanism is similar to in-context learning in large language models, but applied to robotics.

According to the company, the model can chain two prompts to form longer sequences, use demonstrations from simulation environments, and partially mimic human hand movements. These abilities are claimed to emerge spontaneously from over eight months of pretraining on interaction data, rather than being explicitly trained.

Performance Data: From 59% to 83%

Across ten tests, such as opening a jar or taking money from a wallet, GEN-1.5 achieved an average success rate of 59%. When using five minutes of data for ten training steps, the success rate improved to 83%. These figures come from the company's report, but no specific task list or test conditions were provided.

Notably, the 59% baseline success rate indicates that even without training, the model succeeds on most simple tasks, but still fails nearly half the time. While the 83% fine-tuned performance is a significant improvement, the relatively small training steps and data volume may suggest the tasks themselves are simple.

Claimed Generality: First or Continuation?

Generalist claims to be the first team to achieve in-context learning across a wide range of tasks, while other research teams have only demonstrated similar capabilities on a few task types. This claim should be viewed cautiously, as the definition of 'wide range' is vague, and the company has not provided direct comparisons with other methods.

Despite the company's emphasis on generality, the demonstrated tasks are all simple, short-duration operations, and all results come from internal company sources without third-party independent verification. Therefore, this 'first' claim remains a unilateral statement rather than a confirmed fact.

Limitations and Uncertainties

Currently, GEN-1.5's demonstration tasks are simple and short, and whether it can scale to complex, long-horizon tasks remains unclear. The company has not provided analysis of failure cases or explained the model's robustness in the real world.

All performance data comes from the company itself and lacks independent verification, so there is a possibility of selective reporting or overfitting to demonstrations. Additionally, key details such as the model's sensitivity to demonstration quality and its reliance on simulation-to-real transfer have not been disclosed.

Credibility boundary

This report is based on secondary reporting from THE DECODER, and all performance data and capability claims come from Generalist AI, without independent verification. Therefore, these data should be regarded as company claims rather than confirmed facts.

Insight takeaway

GEN-1.5 demonstrates the potential of in-context learning in robotics, but the 59% baseline and 83% fine-tuned success rates are self-reported, and the tasks are simple with no independent verification. Its 'generality' claim requires more evidence, and its actual capabilities should be viewed cautiously until independently replicated.

Primary report

量子位

Primary source

Same-event coverage

Also covered by 1 sources