首页 > AI前沿 > Nous: Learning and Certifying Memory Decisions Before Source Calibration

Nous: Learning and Certifying Memory Decisions Before Source Calibration

arXiv机器学习 2026-10-02 12:00 6 阅读 查看原文

Belief-based agent memory needs reliable decisions about current state, yet its evidence may be noisy, copied, or stale.

Must a memory calibrate its sources before it can improve its decisions? We separate learning, calibration, and revision certification.

On one four-model hidden Markov family

On one four-model hidden Markov family, learning an unknown Bayes decision requires Theta(l^-2) records and certifying its improvement over an informative incumbent takes O(l^-2) fresh records from the same observation law, while fixed-precision source estimation requires Theta(l^-4) as persistence l vanishes.

Thus learning and certifying useful decisions can require quadratically fewer records than source calibration.

A broader model class

A broader model class retains the decision rate and source lower bound.

Under an unknown identity-plus-background report channel

Under an unknown identity-plus-background report channel, we characterize the sharp identified interval for policy improvement and derive a finite-sample certificate using observable witness regions, without pure-class anchors.

A robustness extension

A robustness extension tolerates bounded history-dependent misspecification and conditional copying; split-trained witnesses apply to arbitrary history spaces with explicit power conditions.

Integration and testing

We integrate policy-bound receipts with Nous Dimensions and test 45,000 held-out mutable-state histories and 9,000 episodes in three external MiniGrid memory environments with an introduced noisy-report interface.

The new certificate accepts 9/9 improvements over a constant incumbent and 4/9 over last-write-wins, versus none for the earlier certificate in MiniGrid.

Strong established inference baselines remain competitive or better.

The result is a statistical account of when memory decisions can be learned and justified without recovering source reliability, not a universally superior memory algorithm.