首页 > AI前沿 > Reinforcement learning with prediction-based rewards

Reinforcement learning with prediction-based rewards

OpenAI 2018-10-31 15:00 1 阅读 查看原文

突破性进展:随机网络蒸馏(RND)推动智能体探索

We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments through curiosity, which for the first time exceeds average human performance on Montezuma’s Revenge.