首页 > AI前沿 > FORGE: Verification-Gated Behavioral Repair for Generative Language Models

FORGE: Verification-Gated Behavioral Repair for Generative Language Models

arXiv自然语言 2026-10-04 20:55 5 阅读 查看原文

Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs.

A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally, provide correctness guarantees.

Existing approaches address this only partially: gradient-based fine-tuning lacks per-instance guarantees and becomes unstable with few defect samples; model editing assumes explicit knowledge replacement rather than behavioral correction; and constraint-based repair is largely restricted to discriminative models with unique target outputs.

We present FORGE

We present FORGE, a framework for targeted behavioral repair of generative language models that separates defect localization, weight editing, and behavioral verification into independent stages.

Its core is a repair abstraction that converts localized defective generation into explicit optimization objectives, enabling verification-oriented repair techniques to operate on autoregressive generation.

FORGE is editing-mechanism agnostic: we instantiate it with (1) a constraint-based quadratic optimization method that provides per-sample repair certificates and (2) a null-space projection editor that minimizes interference with the original model distribution, both under the same localization and verification protocol.

Performance on five open-source LLMs

On five open-source LLMs, FORGE consistently achieves larger reductions in bias and toxicity than gradient-based fine-tuning with minor perplexity degradation.

The two backends exhibit complementary performance across architectures, which a lightweight causal probe traces to where toxicity-related signals concentrate.

FORGE also remains effective with only a handful of defective examples, where conventional fine-tuning often oscillates or fails to converge.