APriority78
机器之心
1 sourcesZhejiang University Proposes ProVisE: Let Generative Models 'Draw' Spatial Cognition Instead of Outputting Coordinates
The OmniAI team at Zhejiang University introduces ProVisE, a framework that enables image generation models to answer spatial questions visually rather than through coordinate outputs. By using visual protocols and automated construction, the framework parses generated images into structured predictions, and introduces the SpatialGen-Bench benchmark for evaluation. This approach aims to assess spatial cognition more naturally.