🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Agents · Memory

When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory

First page
When to Remember, When to Abstain: Category-Conditioned Retention for Reliable Agent Memory
The curator’s take

Olukunle Owolabi, Pulkit Gupta and Fei Wang at Meta AI study when an agent's memory pipeline should store an inferred assertion, and propose confidence thresholds conditioned on the assertion's semantic category. The paper is accepted at the NeurIPS 2026 Social Agent Workshop.

Ask this paper

Key points
01

Asymmetry. Across 4,715 candidate assertions from a deployed cold-start memory pipeline on 100 synthetic personas, only 77.9% of value and belief assertions are supported by their source, against 96.2% for every other category.

02

Why a global threshold fails. Values make up 21.5% of candidates but contribute more unsupported assertions than all other categories combined, so one confidence cutoff either admits unsupported value claims or discards well-supported ones.

03

Method. A stricter confidence bar applies only to the values category, fit by 20 repetitions of persona-split five-fold cross-validation.

04

Results. Unsupported retentions fall from 6.2% to 4.0%, about a 36% relative reduction. Against a global threshold at comparable retention (80.3% vs 81.8%), the targeted rule keeps about 13 points more coverage (95% CI 9.8 to 16.0).

05

Takeaway. Whether to write a memory should depend on the type of assertion as well as the model's confidence.

Abstract

Persistent agent memory is only as reliable as its retention decision: an assertion weakly supported by its source can be stored and later reused as established fact. We study whether the retention decision should be governed by a confidence bar conditioned on the semantic category of the assertion rather than by a single global threshold, retaining well-evidenced categories liberally while abstaining more aggressively where inference is unreliable. We evaluate this in a deployed cold-start memory pipeline on 100 synthetic personas. The empirical evaluation is motivated by a sharp reliability asymmetry: across 4{,}715 candidate assertions, only 77.9\% of value and belief assertions are supported by their source, versus 96.2\% for all other categories. A global confidence threshold cannot separate these: it either admits unsupported value claims or discards well-evidenced ones. Conditioning the threshold on category resolves the tradeoff. In repeated held-out evaluation, a stricter bar on values alone reduces unsupported retentions from 6.2\% to 4.0\% (an ${\approx}36\%$ relative reduction, modest but consistent across folds) and, as corroborating evidence, preserves an estimated 13 percentage points more coverage (95\% CI 9.8--16.0) than a global threshold at comparable retention. Our results suggest that reliable retention depends on the type of assertion, not on confidence alone, and that a category-conditioned threshold can act as a simple, effective form of selective prediction at the write boundary.

Every Monday
Get next week’s papers.
Subscribe on Substack