People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings.
Here, we find systematic differences across access modes.
Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity.
Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.