MotionGPT
First page

Paper summary
Generates consecutive human motions from multimodal control signals via LLM instructions.
Ask this paper
01
Motion quantization: Quantizes motion into discrete tokens that LLMs can produce in the same stream as text.
02
Multimodal control: Accepts text, audio, and other control signals as input, producing corresponding human motion outputs.
03
LLM-as-motion-generator: Treats motion generation as a token-prediction task, unifying motion with other LLM capabilities.
04
Animation and VR: Applicable to character animation, VR avatars, and content creation workflows where text-driven motion is valuable.