Spin in clinical trials includes reporting practices that distort the presentation of results. This is particularly critical in medicine, where spin is present in more than 50% of randomized controlled trials that fail to reach statistical significance.
The comparison of primary and reported outcomes is crucial for detecting several types of spin, including outcome switching.
We used 300 pairs of outcomes labeled with semantic similarity to develop a system for automatic detection of outcome switching.
We evaluated baseline text similarity models and open-source LLMs using generated similarity scores and the Youden index to determine the classification threshold.
The proposed approach involves prompt engineering, classification based on token probabilities, and majority voting for the final decision.
The results on the test set of 2,496 examples with an F1 score of 0.78 and an accuracy of 0.90 outperform baseline text similarity models but trail behind fine-tuned versions of BERT.
We used LLMs to generate natural language explanations for the classified instances and manually assessed their quality.