Mainstream large language models rely on a tokenizer to encode text into a token sequence.
Different tokenizers may yield token sequences of substantially different lengths for the same text.
With a fixed model architecture, shorter token sequences correspond to lower inference time.
We Propose a Tokenizer Training Approach Named Counting and Filtering (CNF) and a Text Encoding Algorithm Called Min-Cost Encoding (MCE)
MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost.
CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE.
CNF-MCE Combination Offers Several Advantages Over BPE
- Higher token efficiency
- Greater scalability
- Lower dependency
Experiments and Results
Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers.
With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers.
Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%.
The MCE Algorithm
The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others.
Performance on Language Models
Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks.
These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.