Transformer-based large language models (LLMs) such as RoBERTa represent text using contextual word embeddings (CWEs), which alter the embeddings associated with each token based on surrounding context.
We construct token-wise incremental trajectories by repeatedly recomputing a token's CWE as successive words are added to a sentence, yielding a representation of how contextualized embeddings evolve as the utterance unfolds.
Methodology and Evaluation
We evaluate this approach using garden-path sentences as a test case with characteristic features.
Token-wise trajectories reproduce known features of garden-path processing, including disruption around the critical region, and reliably distinguish garden-path sentences from matched disambiguated controls.
Metric Development
We introduce several metrics for quantifying representational displacement across contextual increments and show that trajectory information can be highly predictive of sentence type.
Findings
We find that ambiguity-related information is recoverable not only from the sentence-level CLS representation but also from ordinary vocabulary tokens, suggesting that utterance-level information is distributed across multiple representational scales.
In exploratory analyses, we find qualitatively similar trajectory structures in other ambiguity- and misdirection-related linguistic phenomena.
Conclusion
Together, these results establish token-wise incremental trajectories as a promising framework for studying utterance-specific meaning construction using LLMs.