HyperAIHyperAI

Command Palette

Search for a command to run...

MEMADAPTER: COUNTERFACTUAL ADAPTATION AGAINST MEMORY-INDUCED SYCOPHANCY

Ruqing Ning Haibo Meng Zhishang Xiang Zerui Chen Jinsong Su Xin Wang Qinggang Zhang

Abstract

Long-term memory enables LLM-based agents to retain and reuse information across tasks and sessions, supporting personalization and long-horizon interactions. However, persistent memories can also induce sycophancy, causing agents to over-align with users’ historical beliefs even when they are inaccurate, outdated, or inconsistent with objective evidence. Existing mitigation methods assume that memory-induced sycophancy originates from biased or incorrect memories and attempt to reduce this risk by filtering such memories at different stages of the memory pipeline. However, in the real world, objective and correct memories can still induce sycophancy, and the same memory can warrant different influence across different contexts. To this end, we propose MemAdapter, a novel framework that adaptively integrates retrieved memories to support objective and reliable reasoning. Specifically, MemAdapter consists of three components: (i) Counterfactual Induction, which leverages counterfactual reasoning to uncover the potential risk of retrieved memories; (ii) Context-Aware Reflection, which calibrates the inferential influence of each retrieved memory in light of the current task via self-reflection; and (iii) Evidence-Based Reasoning, which grounds the final response in appropriate evidence while preserving the legitimate influence of memory. Extensive experiments on three benchmarks demonstrate that MemAdapter consistently improves memory reliability across diverse scenarios. Our code is available at https://github.com/DEEP-JLU/MemAdapter.

One-sentence Summary

Researchers from Jilin University and Xiamen University propose MemAdapter, a framework that adaptively integrates retrieved memories through counterfactual induction, context-aware reflection, and evidence-based reasoning to reduce memory-induced sycophancy in LLM-based agents, addressing the limitations of prior filtering methods and improving memory reliability across three benchmarks.

Key Contributions

  • Identifies a post-retrieval failure mode in long-term memory agents where even accurate and relevant memories can be used beyond what the current task justifies, and introduces MemAdapter, a post-retrieval adaptation framework that dynamically adjusts how retrieved memories participate in reasoning without modifying upstream memory systems.
  • MemAdapter uses Counterfactual Induction to infer conditional-use rules and uncover potential memory risks, Context-Aware Reflection to calibrate each memory’s inferential influence for the current task and evidence, and Evidence-Based Reasoning to ground final responses while preserving legitimate memory influence.
  • Experiments on MemSyco-Bench, Persist-Bench, and MemTrapBench demonstrate consistent improvements in memory-use reliability across different memory systems and backbone models.

Introduction

Long-term memory has become a core component of LLM-based agents, enabling them to retain and reuse user information across tasks and sessions for personalization and long-horizon interaction. However, this same memory can reintroduce past user beliefs, preferences, and judgments into later reasoning, creating memory-induced sycophancy where agents over-align with outdated or inaccurate views. Prior mitigation methods mostly intervene before reasoning at the extraction, retrieval, or organization stages, treating the problem as biased or incorrect memory content. The authors argue that even accurate and task-relevant memories can be misused depending on the reasoning context, making appropriate memory influence a relational property rather than an intrinsic one. To address this gap, they propose MemAdapter, a post-retrieval adaptation framework that uses counterfactual induction, context-aware reflection, and evidence-based reasoning to regulate how retrieved memories influence final answers while preserving their legitimate use.

Method

The authors propose MemAdapter, a framework designed to mitigate memory-induced sycophancy by carefully calibrating the influence of retrieved memories during reasoning. As shown in the figure below, the architecture comprises three core modules: Counterfactual Induction, Context-Aware Reflection, and Evidence-Based Reasoning.

The first module, Counterfactual Induction, addresses the conditional nature of memory-induced risks. Rather than evaluating a memory in isolation, the authors formulate risk identification as a boundary-discovery problem. For each retrieved memory mim_imi​, the model holds the memory content fixed and systematically varies plausible downstream task conditions to construct a counterfactual task space, denoted as Hi={(tij,hij)}j=1Ni\mathcal{H}_i = \{(t_{ij}, h_{ij})\}_{j=1}^{N_i}Hi​={(tij​,hij​)}j=1Ni​​, where tijt_{ij}tij​ represents a task type and hijh_{ij}hij​ instantiates a concrete request and evidence configuration. The model then assesses the appropriate contribution vijv_{ij}vij​ of the memory in each setting, forming the set Vi={(hij,vij)}j=1Ni\mathcal{V}_i = \{(h_{ij}, v_{ij})\}_{j=1}^{N_i}Vi​={(hij​,vij​)}j=1Ni​​. By comparing these setting-specific contributions and grouping settings with equivalent impacts, the module induces a cross-context risk characterization Bi={(air,vir)}r=1RiB_i = \{(a_{ir}, v_{ir})\}_{r=1}^{R_i}Bi​={(air​,vir​)}r=1Ri​​, which provides a task-agnostic boundary of potential risks for subsequent calibration.

The second module, Context-Aware Reflection, bridges the gap between cross-context boundaries and the specific downstream task. Given the current request qqq, available evidence ccc, retrieved memories M\mathcal{M}M, and the conditional boundaries B\mathcal{B}B, the model performs self-reflection to reassess each memory in context. It examines each mim_imi​ alongside its boundary BiB_iBi​ to account for updates, qualifications, and conflicts among the retrieved set. This process yields a task-conditioned representation U={ui}i=1n\mathcal{U} = \{u_i\}_{i=1}^nU={ui​}i=1n​, where each uiu_iui​ serves as a natural-language instruction detailing exactly what the memory may support, which parts of the reasoning it may affect, the extent of its influence, and the conclusions it cannot justify.

The final module, Evidence-Based Reasoning, ensures that the calibrated constraints are strictly respected during answer generation. It prevents useful personalization information from being implicitly promoted into factual evidence. Given the inputs q,c,Mq, c, \mathcal{M}q,c,M, and the influence instructions U\mathcal{U}U, the system jointly generates the final answer yyy and an internal answer-support trace L\mathcal{L}L according to the function (y,L)=REBR(q,c,M,U)(y, \mathcal{L}) = \mathcal{R}_{\mathrm{EBR}}(q, c, \mathcal{M}, \mathcal{U})(y,L)=REBR​(q,c,M,U). The support trace L\mathcal{L}L explicitly associates each major response span with its role, supporting source, and applicable memory-influence instruction. Before returning the final response yyy to the user, the framework verifies that every recorded source is present, supports the corresponding span, and that all memory usage remains within the scope defined by its instruction.

Experiment

A preliminary study shows that even accurate, task-relevant memories can induce sycophancy, so memory risk depends on contextual use rather than content alone. The main experiments evaluate MemAdapter across three memory-use benchmarks and five memory systems, finding consistent improvements in appropriate memory use, robustness across heterogeneous memory systems, and advantages over existing intervention methods. Cross-backbone tests further indicate that these benefits generalize beyond the primary backbone, though gains vary by memory system and task dimension. Ablation results attribute much of the improvement to Counterfactual Induction, with additional benefit from Evidence-Based Reasoning.

MemAdapter improves memory-use behavior over direct generation and existing interventions across MemSyco-Bench, PersistBench, and MemTrapBench. Gains are consistent for MemSyco-Bench dimensions and MemTrapBench scores across evaluated memory systems, while sycophancy failure also decreases on PersistBench. Some PersistBench metrics, particularly beneficial-memory failure, show memory-system-dependent trade-offs. MemAdapter achieves the best When to Use Memory and How to Use Memory accuracy among compared interventions across all evaluated memory systems. Compared with direct generation, MemAdapter improves both MemTrapBench scores across all evaluated memory systems and reduces sycophancy failure rates on PersistBench. In the NaiveRAG block, MemAdapter improves most metrics but slightly increases beneficial-memory failure rate, illustrating a memory-system-dependent trade-off.

MemAdapter improves overall memory-use performance over direct generation across both GPT-5.6-sol and Qwen3-8B backbones. The gains vary by memory system and task dimension, with Qwen3-8B showing especially clear improvements in memory-evidence conflict handling and valid memory selection. Personalized memory use improves consistently with Qwen3-8B but shows mixed changes with GPT-5.6-sol. MemAdapter raises average accuracy relative to direct generation across the evaluated memory systems and backbones. Under Qwen3-8B, gains are particularly pronounced for memory-evidence conflict handling and valid memory selection, while GPT-5.6-sol shows more moderate or mixed changes in some systems. Personalized memory use improves across evaluated memory systems with Qwen3-8B, whereas GPT-5.6-sol exhibits mixed changes, including declines for some memory systems.

Across both tested memory systems, MemAdapter progressively improves memory-use accuracy over direct generation as more components are added. Counterfactual Induction produces broad gains across all task-specific metrics, and the final Evidence-Based Reasoning stage adds further improvement, especially in valid memory selection and average accuracy. The full configuration achieves the highest reported accuracy within each represented memory system. Counterfactual Induction yields consistent improvements across every task-specific metric and substantially raises average accuracy over direct generation for both memory systems. Adding Evidence-Based Reasoning provides further gains, with the full configuration achieving the best average performance and a notable increase in valid memory selection.

The evaluation spans MemSyco-Bench, PersistBench, and MemTrapBench, comparing MemAdapter with direct generation and existing interventions across multiple memory systems and backbones. MemAdapter consistently improves memory-use accuracy, especially for When to Use Memory, How to Use Memory, conflict handling, and valid memory selection, while also reducing sycophancy failure, though some PersistBench beneficial-memory metrics show memory-system-dependent trade-offs. Ablations indicate that Counterfactual Induction provides broad gains and Evidence-Based Reasoning adds further improvements, with the full configuration achieving the best performance.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp