HyperAIHyperAI

Command Palette

Search for a command to run...

Falsche Grenzen: Diagnose und Eindämmung von Co-Cheating bei selbst-evolvierenden Suchagenten

Zusammenfassung

Selbst-evolvierende Suchagenten können ihre eigenen Trainingscurricula erstellen, indem sie gemeinsam einen Proposer, der Fragen generiert, und einen Solver, der sie beantwortet, optimieren. Dieser geschlossene Kreislauf führt zu einem Fehlermodus, den wir Co-Cheating nennen: Proposer und Solver stimmen zunehmend in geteilten Fehlern überein, sodass die interne Belohnung steigt, ohne dass die externe Korrektheit entsprechend zunimmt. Eine Post-hoc-Referenzprüfung anhand von Quellenbelegen zeigt, dass Co-Cheating über aufeinanderfolgende Runden der Selbst-Evolution zunehmend schwerwiegender wird, wobei die Korrektheit der Pseudo-Labels stagniert oder abnimmt, obwohl sich das Trainingssignal innerhalb der Schleife verbessert. Die direkteste Eindämmung besteht darin, jeden Vorschlag vor dem Training zu verifizieren. Wir führen daher Multi-Sample Verification (MSV) ein, die dasselbe in der Selbst-Evolution verwendete Modell dreimal mit der Quelle und dreimal ohne sie abfragt, um über die Aufgabenzulassung zu entscheiden und unzuverlässige Pseudo-Labels zu ersetzen. MSV reduziert die falsche Übereinstimmung teilweise, hinterlässt jedoch erhebliches verbleibendes Co-Cheating und erfordert sechs zusätzliche Labeler-Generierungen für jeden Kandidaten. Diese Einschränkungen motivieren CrossFit, unsere Hauptmethode. Es partitioniert die Quelldokumente des Proposers in die Gruppen A und B: Fragen, die aus A generiert werden, werden von einem Hilfs-Solver bewertet, der nur auf B trainiert wurde, und umgekehrt. Die resultierende Cross-Fitting-Übereinstimmung bestimmt die Proposer-Belohnung und verhindert, dass ein Pseudo-Label aus derselben Quelle direkt über den Feedback-Solver reproduziert wird, während die Aktualisierungsregel des ursprünglichen Solvers unverändert bleibt. Wir evaluieren beide Interventionen, indem wir die vollständige Selbst-Evolutions-Schleife mit Qwen3.5-4B und Qwen3.5-9B erneut durchlaufen. Nach der Selbst-Evolution reduziert MSV die Masse falscher Übereinstimmung von 6,1 % auf 5,7 % bei Qwen3.5-4B und von 8,8 % auf 7,2 % bei Qwen3.5-9B, während CrossFit sie auf 3,0 % bzw. 3,7 % reduziert. Das erneute Abspielen identischer Vorschläge mit quellenausgeschlossenem Feedback reduziert die falsche Übereinstimmung weiter auf 0,4 % und 0,1 % und isoliert die Feedback-Abstammung von Veränderungen im generierten Curriculum. Über sieben nachgelagerte Such-Benchmarks hinweg verbessert CrossFit die durchschnittliche Leistung gegenüber der standardmäßigen gekoppelten Selbst-Evolution um 8,8 bzw. 8,4 Punkte und gegenüber Search-R1 um 8,7 bzw. 7,8 Punkte bei 4B bzw. 9B.

One-sentence Summary

Researchers from Rutgers University; University of California, San Diego; University of Michigan; McGill University; and King Fahd University of Petroleum and Minerals diagnose co-cheating in self-evolving search agents and propose CrossFit, a cross-fitted verification scheme that scores proposals with an auxiliary solver trained on disjoint source partitions, reducing false-agreement mass to 3.0% on Qwen3.5-4B and 3.7% on Qwen3.5-9B while improving average performance over standard coupled self-evolution on seven downstream search benchmarks by 8.8 and 8.4 points, respectively.

Key Contributions

  • The paper identifies and empirically characterizes a failure mode called co-cheating in self-evolving search agents, where the proposer and solver increasingly agree on shared errors so internal reward improves while source-audited pseudo-label correctness stagnates or declines over successive rounds.
  • It proposes multi-sample verification (MSV), which queries the same model three times with source evidence and three times without it to decide task admission and replace unreliable pseudo-labels; MSV lowers false-agreement mass from 6.1% and 8.8% to 5.7% and 7.2% on Qwen3.5-4B and Qwen3.5-9B but leaves substantial residual co-cheating and requires six labeler generations per candidate.
  • It proposes CrossFit as the main method, partitioning source documents into groups A and B so questions generated from A are scored by an auxiliary solver trained only on B, and vice versa; CrossFit reduces false-agreement mass to 3.0% and 3.7% and improves seven-benchmark average performance over coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.

Introduction

Search-augmented language models improve answering by interleaving reasoning with retrieval or browser actions, but self-evolving proposer-solver systems create their own training questions and pseudo-labels, using solver agreement as a reward signal. The authors show that this setup can produce “co-cheating”: an incorrect pseudo-label teaches the solver to repeat the same error, and the proposer is then rewarded for generating questions that reinforce that false agreement. Multi-sample verification reduces this problem only partially and adds inference cost. The authors’ main contribution is CrossFit, which splits proposer source documents into groups A and B, trains auxiliary scoring solvers on complementary groups, and uses cross-fitted agreement to shape proposer rewards, preventing a same-source error from directly reinforcing itself. In experiments with Qwen3.5-4B and Qwen3.5-9B, CrossFit lowers false-agreement mass and improves seven-benchmark question-answering accuracy over coupled self-evolution and Search-R1.

Method

The authors introduce two key interventions within the self-evolution loop to mitigate error reinforcement: Multi-Sample Verification (MSV) and Cross-Fitted Proposer Feedback (CrossFit).

First, MSV acts as an admission-time test to ensure a proposed question yields a stable answer independent of the proposer's draft. Given a source document xxx and a question qqq, the model MMM generates three source-aware answers aisrc∼M(⋅∣x,q)a_i^{\text{src}} \sim M(\cdot \mid x, q)aisrc​∼M(⋅∣x,q) and three source-blind answers aiblind∼M(⋅∣q)a_i^{\text{blind}} \sim M(\cdot \mid q)aiblind​∼M(⋅∣q) for i∈{1,2,3}i \in \{1, 2, 3\}i∈{1,2,3}. A majority function Maj\text{Maj}Maj returns an answer if at least two samples agree under an answer matcher ≃\simeq≃, and ∅\emptyset∅ otherwise. Defining yv=Maj(a1:3v)y^v = \text{Maj}(a_{1:3}^v)yv=Maj(a1:3v​) for v∈{src,blind}v \in \{\text{src}, \text{blind}\}v∈{src,blind}, the admission criterion is formulated as:

IMSV=1[ysrc≠∅∧yblind≠∅∧ysrc≃yblind].I_{\text{MSV}} = \mathbf{1} \left[ y^{\text{src}} \neq \emptyset \land y^{\text{blind}} \neq \emptyset \land y^{\text{src}} \simeq y^{\text{blind}} \right].IMSV​=1[ysrc=∅∧yblind=∅∧ysrc≃yblind].

When IMSV=1I_{\text{MSV}} = 1IMSV​=1, the compatible majority replaces the draft as the training label; otherwise, the task is rejected. This step improves the quality of supervision entering the training phase.

While MSV refines the training labels, CrossFit prevents this supervision from being directly recycled into the proposer's reward. The overall pipeline for this cross-fitted feedback mechanism is detailed in the framework diagram below.

In this architecture, the proposer first generates questions and pseudo-labels from source documents. The authors then partition the source documents into two distinct groups, fold 0 and fold 1. This split is performed at the source level to prevent related examples from the same document from being placed on both sides, which would preserve the data reuse path they aim to eliminate.

The system maintains two auxiliary feedback solvers corresponding to these folds. During round rrr, one auxiliary solver trains exclusively on admitted questions from fold 0, while the other trains only on fold 1. In the subsequent round r+1r+1r+1, their evaluation roles are crossed: questions originating from fold 0 are scored by the solver trained on fold 1, and questions from fold 1 are scored by the solver trained on fold 0. This ensures that the solver evaluating a specific question has not been trained on pseudo-labels derived from that question's source.

The feedback rule calculates the reward based on the complementary solver's responses. Let hhh denote the source fold, Sr,1−hS_{r, 1-h}Sr,1−h​ the auxiliary solver trained on the complementary fold, and y~\tilde{y}y~​ the adopted label. The proposer receives a reward based on five rollouts z1,…,z5z_1, \ldots, z_5z1​,…,z5​:

RP(q)=f(∑j=151[zj≃y~]),zj∼Sr,1−h(⋅∣q).R_P(q) = f \left( \sum_{j=1}^5 \mathbf{1} [ z_j \simeq \tilde{y} ] \right), \qquad z_j \sim S_{r, 1-h}(\cdot \mid q).RP​(q)=f(j=1∑5​1[zj​≃y~​]),zj​∼Sr,1−h​(⋅∣q).

Here, the function f(k)=(5−k)/4f(k) = (5 - k) / 4f(k)=(5−k)/4 for 0<k<50 < k < 50<k<5 (and zero otherwise) serves as the frontier reward. Consequently, the original training objective is preserved, as questions still receive credit based on their perceived difficulty to the solver, but the direct self-reinforcing error path is broken.

Distinct from the auxiliary solvers, the main solver is not split. It continues to train on all admitted questions from both folds, utilizing the refined feedback to shape the proposer's curriculum for the next round without inheriting the localized source-derived errors.

Experiment

The experiments evaluate self-evolution on seven open-domain question answering benchmarks using Qwen3.5-4B and Qwen3.5-9B, comparing the coupled Dr. Zero loop with multi-sample verification and source-excluded CrossFit feedback over three rounds. Audits of the coupled loop show that proposer-solver agreement becomes increasingly optimistic as false agreement accumulates while correctness does not, revealing a co-cheating dynamic. CrossFit markedly improves downstream search over Dr. Zero and Search-R1, especially on multi-hop tasks, whereas verification alone adds little, indicating that feedback provenance matters more than pseudo-label quality alone. Ablations with fixed question banks and source-level splits confirm that excluding the evaluated source from the feedback solver, rather than evaluator duplication, partitioning, or extra updates, prevents shared same-source errors and accounts for most of the learned search improvement.

Cross-fitted feedback improves downstream search performance at both Qwen3.5 scales, with every evaluated benchmark gaining over the comparison methods. Gains are largest on multi-hop tasks, while verification alone provides only a small average improvement and adds little beyond cross-fitting alone. The results point to feedback provenance as the main factor in the improved search policy. CrossFit improves over Dr. Zero and Search-R1 at both 4B and 9B scales, with all benchmarks gaining. Multi-hop benchmarks show larger average gains than single-hop datasets. MSV alone adds less than one point over Dr. Zero, and combining MSV with CrossFit yields only a small additional improvement.

The experiments assess cross-fitted feedback in downstream search using Qwen3.5 at 4B and 9B scales, comparing CrossFit against Dr. Zero and Search-R1. CrossFit consistently improves performance on all benchmarks at both scales, with the largest gains on multi-hop tasks, indicating that feedback provenance is the main driver of the improved search policy. Verification alone provides only marginal average improvement over Dr. Zero and adds little beyond cross-fitting alone, suggesting its contribution is limited.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp