We introduce a calibration-first framework that produces supervision scores without access to ground-truth labels or a shared annotation space.
Our framework aligns subset-specific scorers using a synthetic ordinal reference space before fusion.
This reference space is constructed from ordered calibration features that represent the latent concept, providing a common scale on which otherwise incomparable scorer outputs can be aligned.
Because our calibration procedure uses the reference space rather than training samples, it is independent of the training set's empirical distribution.
Results
Across three benchmark datasets, our framework consistently outperforms uncalibrated averaging and achieves higher primary-metric point estimates on the evaluation metrics than the best individual scorer.
Performance relative to sample-dependent baselines varies by domain, with absolute differences below 0.02 on Ames Housing and below 0.01 on Breast Cancer Wisconsin and Wine Quality.
After Bonferroni correction, differences remain significant for all three comparisons on Ames Housing and one on Breast Cancer Wisconsin.
Additionally, we show that using fewer calibration levels per feature can closely approximate higher-resolution results at substantially lower computational cost.
Conclusion
Together, these results support our framework as a viable approach to construct supervision scores when neither ground-truth labels nor a shared annotation space is available.