For multimodality, the model can process documents over 200 pages and videos over 100 hours. The team introduced the RecreationBench benchmark, requiring the model to rebuild applications without source code, only by observing clicks and keyboard interactions, covering Ubuntu, macOS, Windows, Android, and Web. They also released the Qwen-MM-Plugins extension library, adding image/video processing, visual tool use, and multimodal memory to existing agent systems.
The team's published benchmarks show the model approaching or surpassing Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol in several categories. PaperBench scored 93, the highest among comparisons, while TerminalBench 2.1 scored 86.6, below GPT-5.6 Sol's 88.8. These are internal runs, with independent verification pending.
The team attributes long-task capability to expanded reinforcement learning training environments, covering multi-day workflows, nested directory structures, and various agent frameworks. The internal score index rose from 0.474 to 0.725, with peak performance at around 4,000 environments, then slightly declining.