Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment.
Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student.
The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows.
We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student's training target.
On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors.
With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child.
With the evaluated children's forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels.
After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels.
Their accuracy largely follows label quality.
The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers.
Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets.
Code is available at https://github.com/js-lee-AI/directional-verification.