Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples.
We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures.
We formalize Policy Iteration with Human Feedback (PIHF)
PIHF makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP.
Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies.
Performance Improvements
Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively.
With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores.
Low-Data Rare-Disease Diagnosis Setting
In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points).
Efficient Inference-Time Scaling
These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models.
Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.