首页 > AI前沿 > Improving Model Safety Behavior with Rule-Based Rewards

Improving Model Safety Behavior with Rule-Based Rewards

OpenAI 2024-07-24 17:00 1 阅读 查看原文

We’ve developed and applied a new method leveraging Rule-Based Rewards (RBRs) that aligns models to behave safely without extensive human data collection.