Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge.
Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the functional heterogeneity of commentary.
We introduce \textsc{GameCommBench}, a unified benchmark spanning board games, sports, and esports, with commentary aligned to heterogeneous game contexts and annotated by commentary type.
We further propose Type-Aware Commentary Evaluation (TACE), a structured framework for evaluating different types of commentary.
We then validate TACE for reliability and human agreement, and use it to benchmark representative AI commentators.
Results reveal non-uniform capability profiles, with live observation and strategic analysis emerging as major bottlenecks.
Together, \textsc{GameCommBench} and TACE provide a diagnostic foundation for comparable and interpretable AI-GGC evaluation.