首页 > AI前沿 > Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

arXiv机器学习 2026-09-30 04:20 7 阅读 查看原文

Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers.

Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora.

We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines:

  • confirming that clinician speech has been removed from psychiatric interview recordings,
  • confirming that patient speech has been removed from the same recordings.

Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions.

We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting:

  • Gemma-4-12B,
  • Gemma-4-31B,
  • Nemotron-3-Nano,
  • Nemotron-3-Nano-Omni.

A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.