HyperAIHyperAI

Command Palette

Search for a command to run...

Fausses frontières : diagnostic et atténuation du co-cheating dans les agents de recherche auto-évolutifs

Résumé

Les agents de recherche auto-évolutifs peuvent construire leur propre curriculum d’entraînement en optimisant conjointement un proposeur qui génère des questions et un solveur qui y répond. Cette boucle fermée introduit un mode de défaillance que nous appelons co-cheating : le proposeur et le solveur convergent de plus en plus vers des erreurs partagées, de sorte que la récompense interne s’améliore sans augmentation correspondante de la justesse externe. Un audit de référence post-hoc réalisé sur les sources montre que le co-cheating devient de plus en plus marqué au fil des cycles successifs d’auto-évolution, la justesse des pseudo-étiquettes stagnant ou diminuant alors même que le signal d’entraînement en boucle s’améliore. La parade la plus directe consiste à vérifier chaque proposition avant l’entraînement. Nous présentons donc la vérification multi-échantillons (MSV), qui interroge le même modèle que celui utilisé lors de l’auto-évolution trois fois avec la source et trois fois sans elle afin de décider de l’admission de la tâche et de remplacer les pseudo-étiquettes peu fiables. La MSV réduit partiellement l’accord erroné, mais laisse un co-cheating résiduel substantiel et exige six générations d’étiquetage supplémentaires pour chaque candidat. Ces limites motivent CrossFit, notre méthode principale. Elle répartit les documents sources du proposeur en groupes A et B : les questions générées à partir de A sont évaluées par un solveur auxiliaire entraîné uniquement sur B, et inversement. L’accord ainsi obtenu par ajustement croisé détermine la récompense du proposeur, empêchant qu’une pseudo-étiquette issue de la même source soit directement reproduite par le solveur de rétroaction, tout en laissant inchangée la règle de mise à jour du solveur d’origine. Nous évaluons ces deux interventions en réexécutant la boucle complète d’auto-évolution avec Qwen3.5-4B et Qwen3.5-9B. Après auto-évolution, la MSV réduit la masse d’accords erronés de 6,1 % à 5,7 % pour Qwen3.5-4B et de 8,8 % à 7,2 % pour Qwen3.5-9B, tandis que CrossFit la réduit à 3,0 % et 3,7 %, respectivement. La réexécution des mêmes propositions avec une rétroaction excluant la source réduit encore l’accord erroné à 0,4 % et 0,1 %, ce qui isole l’origine de la rétroaction des changements dans le curriculum généré. Sur sept bancs d’essai de recherche en aval, CrossFit améliore la performance moyenne de 8,8 et 8,4 points par rapport à l’auto-évolution couplée standard, et de 8,7 et 7,8 points par rapport à Search-R1, à 4B et 9B, respectivement.

One-sentence Summary

Researchers from Rutgers University; University of California, San Diego; University of Michigan; McGill University; and King Fahd University of Petroleum and Minerals diagnose co-cheating in self-evolving search agents and propose CrossFit, a cross-fitted verification scheme that scores proposals with an auxiliary solver trained on disjoint source partitions, reducing false-agreement mass to 3.0% on Qwen3.5-4B and 3.7% on Qwen3.5-9B while improving average performance over standard coupled self-evolution on seven downstream search benchmarks by 8.8 and 8.4 points, respectively.

Key Contributions

  • The paper identifies and empirically characterizes a failure mode called co-cheating in self-evolving search agents, where the proposer and solver increasingly agree on shared errors so internal reward improves while source-audited pseudo-label correctness stagnates or declines over successive rounds.
  • It proposes multi-sample verification (MSV), which queries the same model three times with source evidence and three times without it to decide task admission and replace unreliable pseudo-labels; MSV lowers false-agreement mass from 6.1% and 8.8% to 5.7% and 7.2% on Qwen3.5-4B and Qwen3.5-9B but leaves substantial residual co-cheating and requires six labeler generations per candidate.
  • It proposes CrossFit as the main method, partitioning source documents into groups A and B so questions generated from A are scored by an auxiliary solver trained only on B, and vice versa; CrossFit reduces false-agreement mass to 3.0% and 3.7% and improves seven-benchmark average performance over coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.

Introduction

Search-augmented language models improve answering by interleaving reasoning with retrieval or browser actions, but self-evolving proposer-solver systems create their own training questions and pseudo-labels, using solver agreement as a reward signal. The authors show that this setup can produce “co-cheating”: an incorrect pseudo-label teaches the solver to repeat the same error, and the proposer is then rewarded for generating questions that reinforce that false agreement. Multi-sample verification reduces this problem only partially and adds inference cost. The authors’ main contribution is CrossFit, which splits proposer source documents into groups A and B, trains auxiliary scoring solvers on complementary groups, and uses cross-fitted agreement to shape proposer rewards, preventing a same-source error from directly reinforcing itself. In experiments with Qwen3.5-4B and Qwen3.5-9B, CrossFit lowers false-agreement mass and improves seven-benchmark question-answering accuracy over coupled self-evolution and Search-R1.

Method

The authors introduce two key interventions within the self-evolution loop to mitigate error reinforcement: Multi-Sample Verification (MSV) and Cross-Fitted Proposer Feedback (CrossFit).

First, MSV acts as an admission-time test to ensure a proposed question yields a stable answer independent of the proposer's draft. Given a source document xxx and a question qqq, the model MMM generates three source-aware answers aisrc∼M(⋅∣x,q)a_i^{\text{src}} \sim M(\cdot \mid x, q)aisrc​∼M(⋅∣x,q) and three source-blind answers aiblind∼M(⋅∣q)a_i^{\text{blind}} \sim M(\cdot \mid q)aiblind​∼M(⋅∣q) for i∈{1,2,3}i \in \{1, 2, 3\}i∈{1,2,3}. A majority function Maj\text{Maj}Maj returns an answer if at least two samples agree under an answer matcher ≃\simeq≃, and ∅\emptyset∅ otherwise. Defining yv=Maj(a1:3v)y^v = \text{Maj}(a_{1:3}^v)yv=Maj(a1:3v​) for v∈{src,blind}v \in \{\text{src}, \text{blind}\}v∈{src,blind}, the admission criterion is formulated as:

IMSV=1[ysrc≠∅∧yblind≠∅∧ysrc≃yblind].I_{\text{MSV}} = \mathbf{1} \left[ y^{\text{src}} \neq \emptyset \land y^{\text{blind}} \neq \emptyset \land y^{\text{src}} \simeq y^{\text{blind}} \right].IMSV​=1[ysrc=∅∧yblind=∅∧ysrc≃yblind].

When IMSV=1I_{\text{MSV}} = 1IMSV​=1, the compatible majority replaces the draft as the training label; otherwise, the task is rejected. This step improves the quality of supervision entering the training phase.

While MSV refines the training labels, CrossFit prevents this supervision from being directly recycled into the proposer's reward. The overall pipeline for this cross-fitted feedback mechanism is detailed in the framework diagram below.

In this architecture, the proposer first generates questions and pseudo-labels from source documents. The authors then partition the source documents into two distinct groups, fold 0 and fold 1. This split is performed at the source level to prevent related examples from the same document from being placed on both sides, which would preserve the data reuse path they aim to eliminate.

The system maintains two auxiliary feedback solvers corresponding to these folds. During round rrr, one auxiliary solver trains exclusively on admitted questions from fold 0, while the other trains only on fold 1. In the subsequent round r+1r+1r+1, their evaluation roles are crossed: questions originating from fold 0 are scored by the solver trained on fold 1, and questions from fold 1 are scored by the solver trained on fold 0. This ensures that the solver evaluating a specific question has not been trained on pseudo-labels derived from that question's source.

The feedback rule calculates the reward based on the complementary solver's responses. Let hhh denote the source fold, Sr,1−hS_{r, 1-h}Sr,1−h​ the auxiliary solver trained on the complementary fold, and y~\tilde{y}y~​ the adopted label. The proposer receives a reward based on five rollouts z1,…,z5z_1, \ldots, z_5z1​,…,z5​:

RP(q)=f(∑j=151[zj≃y~]),zj∼Sr,1−h(⋅∣q).R_P(q) = f \left( \sum_{j=1}^5 \mathbf{1} [ z_j \simeq \tilde{y} ] \right), \qquad z_j \sim S_{r, 1-h}(\cdot \mid q).RP​(q)=f(j=1∑5​1[zj​≃y~​]),zj​∼Sr,1−h​(⋅∣q).

Here, the function f(k)=(5−k)/4f(k) = (5 - k) / 4f(k)=(5−k)/4 for 0<k<50 < k < 50<k<5 (and zero otherwise) serves as the frontier reward. Consequently, the original training objective is preserved, as questions still receive credit based on their perceived difficulty to the solver, but the direct self-reinforcing error path is broken.

Distinct from the auxiliary solvers, the main solver is not split. It continues to train on all admitted questions from both folds, utilizing the refined feedback to shape the proposer's curriculum for the next round without inheriting the localized source-derived errors.

Experiment

The experiments evaluate self-evolution on seven open-domain question answering benchmarks using Qwen3.5-4B and Qwen3.5-9B, comparing the coupled Dr. Zero loop with multi-sample verification and source-excluded CrossFit feedback over three rounds. Audits of the coupled loop show that proposer-solver agreement becomes increasingly optimistic as false agreement accumulates while correctness does not, revealing a co-cheating dynamic. CrossFit markedly improves downstream search over Dr. Zero and Search-R1, especially on multi-hop tasks, whereas verification alone adds little, indicating that feedback provenance matters more than pseudo-label quality alone. Ablations with fixed question banks and source-level splits confirm that excluding the evaluated source from the feedback solver, rather than evaluator duplication, partitioning, or extra updates, prevents shared same-source errors and accounts for most of the learned search improvement.

Cross-fitted feedback improves downstream search performance at both Qwen3.5 scales, with every evaluated benchmark gaining over the comparison methods. Gains are largest on multi-hop tasks, while verification alone provides only a small average improvement and adds little beyond cross-fitting alone. The results point to feedback provenance as the main factor in the improved search policy. CrossFit improves over Dr. Zero and Search-R1 at both 4B and 9B scales, with all benchmarks gaining. Multi-hop benchmarks show larger average gains than single-hop datasets. MSV alone adds less than one point over Dr. Zero, and combining MSV with CrossFit yields only a small additional improvement.

The experiments assess cross-fitted feedback in downstream search using Qwen3.5 at 4B and 9B scales, comparing CrossFit against Dr. Zero and Search-R1. CrossFit consistently improves performance on all benchmarks at both scales, with the largest gains on multi-hop tasks, indicating that feedback provenance is the main driver of the improved search policy. Verification alone provides only marginal average improvement over Dr. Zero and adds little beyond cross-fitting alone, suggesting its contribution is limited.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp