Redpine Science gives models and agents a single access point to a wide range of peer-reviewed literature, queried directly through the Model Context Protocol (MCP) and an API.
This report evaluates Redpine Science on two levels: the relevance of the retrieved chunks, and a model's answer when it has access to Redpine Science compared to web search. Both public and expert-validated benchmarks are used.
Public benchmarks are a widely accepted way to test model development and are comparable across labs, but risk saturation and memorization. To address this, we complement them with an expert-validated question set.
In total, this report presents four evaluations.
On ScholarQABench SciFact
On the public answer-quality benchmark reported here, an agent with Redpine Science answers 94.4% of claims correctly against 87.6% with no retrieval.
On the expert-validated question set
An agent with Redpine Science states 80.1% of the required claims against 70.2% for an agent restricted to web search.
On the 668 queries of a public retrieval benchmark
Whose gold paper Redpine holds, stripped of any model reasoning, Redpine Science places the correct source paper in its top ten results for 83.1% of queries (Recall@10), against 79.3% for the benchmark's creator.
A blinded expert relevance panel
Places Redpine Science's Precision@5 at 75.2% against 39.8% for the PubMed search tool.
We release the expert-validated question set and instructions to reproduce every headline result above, at https://github.com/redpine-ai/benchmarks.