首页 > AI前沿 > Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?

Do Motion Tokenizers for Co-Speech Gesture Generation Encode Gesture Semantics?

arXiv自然语言 2026-09-28 16:50 4 阅读 查看原文

Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation.

It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks.

We probe a reconstruction-trained codebook using 19 co-speech gesture descriptors spanning from raw motion to abstract communicative function.

Results show that geometry and handedness are readily decodable from token embeddings, while motion category is only weakly decoded despite showing systematic differences in discrete code usage.

This gap between reconstruction quality and descriptor decodability suggests that reconstruction objectives alone do not guarantee that gesture semantics are captured, and that evaluating codebooks on such properties can guide the design of more semantic motion tokenizers.