This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation.
Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the issue of the two orthographies of the Belarusian language, noise in the training data, interference from other languages and other misspelling issues common in Belarusian on the internet.
A matched ablation on unfiltered training data shows substantial benefits from filtering for all fine-tuned models, with the LLM-based models gaining roughly twice as much from filtering as the dedicated encoder-decoder MT system.
Supporting the view that for Belarusian MT one of the primary bottlenecks is data quality.