首页 > AI前沿 > Language-Specific Effects of Tokenizer Choice in Multilingual Language Models

Language-Specific Effects of Tokenizer Choice in Multilingual Language Models

arXiv自然语言 2026-10-08 23:29 4 阅读 查看原文

Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others.

We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered.

To this end, we train 123 language models spanning 54 tokenizers.

In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer.

We find that tokenizer choice matters more for languages with less language-model training data:

across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries).

Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data.

Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages:

both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data.

Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language.

We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.