GenCeption's training data is almost entirely synthetic, consisting of only 7,500 video clips. The team combined 800 digital human models with 200 motion capture sequences, rendered in Blender with various backgrounds and viewpoints. Real videos were used only for language-guided segmentation tasks.
According to the paper, GenCeption matches or exceeds existing specialized models on multiple benchmarks: depth estimation is on par with DepthAnything 3, surface normal estimation outperforms NormalCrafter and Lotus-2, 3D pose recognition surpasses Genmo and TRAM, and complex language-guided segmentation matches the combination of Meta's SAM 3 and Gemini 3.5 Flash. Its data volume is only 1/7 to 1/500 that of models like D4RT and VGGT Omega.
Under identical conditions, video generation pre-training also outperforms V-JEPA and VideoMAE V2. The authors infer that this advantage stems from the generative task itself rather than data scale—video generation forces the model to learn useful spatial and motion representations.