Pre-training Large Language Models (LLMs) from scratch at larger scales yields remarkable performance but incurs substantially high training costs.
Depth up-scaling provides an efficient alternative by inserting new layers into a pre-trained LLM, avoiding training from scratch.
However, most existing methods copying or averaging base layers for new layer, which misaligns functionally corresponding neurons, leading to neuron permutation mismatch that harms performance.
To address this issue, we propose Optimal Transport Depth Up-Scaling (OT-DUS), which leverages Optimal Transport (OT) theory to align and fuse functionally corresponding neurons module by module in adjacent base layers for new layer construction.
OT-DUS achieves better overall performance in both general and specialized domains than existing methods for continual pre-training and supervised fine-tuning across different model sizes and model families.
Our analysis of insertion strategies further finds that inserting new layers at higher positions yields not only stronger performance but also improved training efficiency.