Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated.
We propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty.
Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives:
- whether the same idea has appeared before (Exact-Match P-Novelty)
- whether the modified variable or variable combination has been explored before (Variable-level P-Novelty), which reflects the breadth of research-space exploration
- whether the proposed idea is explicitly attributed to external knowledge in the model's reasoning (H-Novelty)
Our evaluation shows that the models achieve relatively similar Valueness and Exact-Match P-Novelty scores, while differing substantially in Variable-level P-Novelty.
H-Novelty is also consistently high among the models for which it can be evaluated.
Notably, further correlation and idea-level performance analyses reveal a strong positive correlation between Variable-level P-Novelty and research performance.