Recent advancements have endowed Large Language Models with impressive general reasoning capabilities.
However, these reasoning models often perform worse than non-reasoning models on personalization tasks.
While some methods use outcome-based RL to improve personalization reasoning, they fail to supervise the reasoning process.
As a result, models may reach correct answers through flawed reasoning chains, limiting further improvement.
To address this
we propose TagPR, a novel framework that adds semantic tags to the reasoning process for step-by-step guidance.
TagPR first automatically generates a structured, tagged dataset for Supervised Fine-Tuning.
It then employs a multi-stage RL process guided by a composite reward signal, which integrates tag-based process supervision with a novel Personalization Reward Model with User Embeddings to achieve fine-grained alignment with user-specific logic.
Experiments and Results
Extensive experiments on public LaMP, LongLaMP, PGraphRAG, and a self-constructed dataset demonstrate that our approach achieves state-of-the-art results,
delivering an average improvement of 32.65% over the base model across all LaMP benchmark tasks.
Conclusion
Our work demonstrates that tag-guided process supervision is an effective approach for personalization reasoning.