Tool retrieval is a critical bottleneck for LLM-based agents operating over large, heterogeneous API ecosystems.
Existing approaches face an inherent trade-off: semantic retrievers are fast but suffer from the semantic-functional gap, while execution-based validation improves precision at the cost of prohibitive latency.
We propose Lookahead-R, a planning-based framework that reformulates tool retrieval as a resource-constrained sequential decision-making problem.
At its core, Lookahead-R introduces a lightweight execution-aware surrogate world model that jointly predicts tool execution success, latency cost, and semantic utility---without invoking real APIs.
This world model drives a cost-sensitive, uncertainty-guided Monte Carlo Tree Search that navigates the tool space under strict budget constraints.
Evaluated on the large-scale ToolBench benchmark, Lookahead-R achieves a superior accuracy-efficiency trade-off across all test scenarios.
On the most challenging I3 split, it attains an NDCG@5 of 91.40%, outperforming the state-of-the-art ToolGen (90.16%) by 1.24%.
Ablation studies confirm that explicit latency modeling is the key discriminative signal for identifying high-quality tools under resource constraints.