Self-supervised fine-tuning refines the embedding space of a pretrained language encoder without labels. However, the commonly used approaches are computationally expensive.
Specifically, contrastive learning-based methods need multiview data and in-batch negative examples, while negative-free approaches require auxiliary graphs/networks.
An interesting question arises: can self-supervised fine-tuning be done without relying on either additional negatives or graph data?
To answer this question, we introduce SilK (Silhouette-guided K-means), which trains on a Cluster Validation Index, an internal measure of cluster quality without using labels.
SilK clusters the corpus and then regresses a simplified silhouette toward a target value.
Each document is compared only against the k cluster centroids, never against other documents, so the method needs no augmentation, no negative pairs and one view per document.
On BERT-base, SilK trains 1.46x faster per epoch than the fastest baseline we evaluate and uses 45.4% less peak GPU memory than the leanest one.
Under frozen-encoder linear probing, SilK stays competitive with the best baselines on three downstream tasks.