首页 > AI前沿 > MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

arXiv机器学习 2026-09-23 12:39 6 阅读 查看原文

Real-world molecular design remains challenging for large language model (LLM)-based agents.

It requires them to interpret design contexts, satisfy multiple constraints, identify infeasible specifications, and reason over multi-step tool outputs.

Existing benchmarks do not capture this complexity, focusing instead on explicit and narrow constraints, only feasible problems, and single-path solutions.

To address this gap, we propose MolDesignBench

a scenario-grounded benchmark that more closely reflects real-world molecular design for evaluating tool-augmented LLM agents.

MolDesignBench comprises 2K generation and optimization instances that combine implicit requirements embedded in design narratives with explicit property and functional-group constraints, including infeasible cases, and require the effective use of 17 specialized chemistry tools.

Experiments across diverse frontier LLMs reveal low success rates--with the best achieving only $\sim43$\%--and frequent failures in implicit-constraint reasoning, infeasibility detection, and tool reasoning.

The corresponding fine-grained failure-mode analysis identifies implicit constraint interpretation and infeasibility detection as the primary bottlenecks, establishing MolDesignBench as a rigorous testbed to guide future research on chemical agents.

The benchmark, tool interface, and evaluation code are publicly available.