首页 > AI前沿 > Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

arXiv机器学习 2026-09-15 23:05 3 阅读 查看原文

Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-specific fine-tuning remain valuable for open-ended scientific reasoning?

We study this in astronomy with a curated QA benchmark from publicly available 2017--2026 Olympiad-style materials.

The free-response subset contains 300 questions, including 204 text-only and 96 image-linked examples.

We compare open-weight and API-served general-purpose, multimodal, and astronomy-specialized models using judge-based correctness and complementary reference metrics.

Strong general-purpose models establish the highest correctness baseline in this testbed, while analyses of metric agreement, judge sensitivity, benchmark composition, and modality reveal variation not captured by a single leaderboard.

These results motivate treating domain specialization as a task- and deployment-dependent property and highlight the role of domain-specific evaluation in determining which models, capabilities, and evaluation criteria are appropriate for scientific workflows.