首页 > AI前沿 > SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children

SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children

arXiv自然语言 2026-08-04 14:53 4 阅读 查看原文

Language is the target of most early intervention for autistic children.

Because the goal and the method change from child to child, the work falls to a teacher who takes one child at a time and judges each scene as it unfolds.

Artificial intelligence is now being brought to that work, yet the benchmarks that reach special education ask what a model knows rather than what it does in front of a child.

Building one is not straightforward, since whether a response is good teaching depends on what the child has just done, so no answer key applies.

The evidence that settles it is visual as much as verbal, since the length of a wait, a shift of gaze, and the child's uptake leave no trace in a transcript.

We introduce SpecialEduBench, which measures pedagogical competence along knowledge, skill, and attitude, with 4,537 knowledge items and with 200 skill items and 68 attitude items built on recorded intervention, the attitude items crossing pressure with monitoring into 192 response cells.

Seven special-education experts wrote, scored, and reviewed the items, and we revised the judge model's instruction against the reference scores they set.

Across eight frontier vision-language models no axis is saturated, since the strongest still fails about a tenth of the honesty cells.

The models converge where the knowledge is factual and separate where the task is situated, and the failures gather where pressure is applied.

We intend the benchmark as an audit to run before deployment and as a starting point for models built for this domain.