首页 > AI前沿 > MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

arXiv自然语言 2026-09-29 12:21 5 阅读 查看原文

Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule.

Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response.

We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories.

MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization.

We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available.

Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting.

Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.