The same interface supports generating continuous motion from text, summarizing action descriptions from continuous trajectories, editing source motion according to text instructions (e.g., changing 'waving' to 'drinking coffee'), continuing future trajectories based on observed prefixes (the example provides 1.05 seconds of action, predicting the next 4.35 seconds), as well as recovering human pose from images, generating descriptions for images or videos, and letting reference motion participate in human image editing. Motion thus can serve as input, output, and cross-modal generation condition.
Experiments cover datasets like HumanML3D, MotionFix, Human3.6M, and MoVid. Radar charts show UniMotion is the only model among comparisons covering all task directions; the paper reports competitive results on Text-to-Motion, Motion-to-Text, motion prediction, motion editing, and visual human pose recovery.
However, the paper also admits that Text-to-Motion distribution metrics are not all optimal, and specialized human pose recovery models still maintain lower error. Under complex prompts like 'a person stomps up some stairs', UniMotion executes continuous leg lifting, weight transfer, and stepping force more completely than MoMask and MotionGPT.