首页 > AI前沿 > AI-Enabled Quality Assurance for Multiple-Choice Assessment Items

AI-Enabled Quality Assurance for Multiple-Choice Assessment Items

arXiv自然语言 2026-10-03 11:45 4 阅读 查看原文

Generating multiple-choice questions is increasingly scalable, but establishing their assessment quality remains difficult.

We present a focused narrative review of automated item-writing flaw detection, revision, psychometric screening, and NLP benchmark auditing.

Database searches, citation retrieval, and nominated sources yield fourteen research reports reviewed in full text.

We distinguish surface checks from content-sensitive judgments and map a 19-criterion rubric to detection methods and reported evidence.

High label-level accuracy often coexists with weak positive case detection, while rubric definitions and reference standards vary.

Revision evidence is mixed, and the associations reported in prior work do not establish the effects of repair.

We propose evaluating quality assurance as a sequence of independently validated decisions, with criterion-specific reporting, calibrated human review, and outcome-based assessment of revisions.