首页 > AI前沿 > From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation

From Answers to Policies: Efficient In-Context Learning System through Emulating Expert Investigation

arXiv自然语言 2026-10-01 12:00 8 阅读 查看原文

Pretrained large language models offer a practical foundation for learning useful behavior from few task-specific examples.

We argue that current prompt and context optimization methods underuse the extensive knowledge and reasoning capabilities of trillion-parameter models. These capabilities can make adaptation more sample-efficient, more compute efficient and at no performance loss when organized around how human experts investigate failures.

We formalize Policy Iteration with Human Feedback (PIHF)

PIHF makes this implicit procedure explicit for LLM agents to execute, and build its automated implementation, PIHF-MCP.

Initialized from clinician feedback on rare-disease diagnosis, PIHF-MCP supplies the expert procedure, testing tools, review and persistent inquiry records to develop reusable task policies.

Performance Improvements

Across general reasoning benchmarks (BIG-Bench Extra Hard, HoVer and LiveBench-Math), PIHF-MCP improved performance of the baseline model by 16.9, 22.2 and 4.7 percentage points, respectively.

With a matched baseline model, development used about 1/5 of the labelled examples and 4% of the task rollouts reported by a previous SOTA in-context optimizer, making it about 9 times faster and 3 times cheaper at comparable or higher scores.

Low-Data Rare-Disease Diagnosis Setting

In a low-data rare-disease diagnosis setting, policies developed from previous SOTA prompt optimizers trailed a previously published PIHF-developed system on every held-out cohort (on average 16 percentage points).

Efficient Inference-Time Scaling

These findings support a route to more efficient inference-time scaling: PIHF-MCP develops reusable policies from a few examples that improve performance on unseen cases and across models.

Because each policy comes from an explicit, recorded investigation, the process also keeps humans in the loop and enables ownership and learning, making it well suited to high-stakes decisions.