Command Palette
Search for a command to run...
EVISKILL : ancrer l'évolution des compétences dans des preuves rejouables
EVISKILL : ancrer l'évolution des compétences dans des preuves rejouables
Yan Zhou Yili Wang Yiwei Dai Qinggang Zhang Xin Wang
Résumé
L'évolution continue des compétences permet aux agents fondés sur des LLM d'accumuler et d'affiner des connaissances procédurales réutilisables à partir de l'expérience d'interaction, sans mettre à jour les paramètres du modèle. Son efficacité dépend de la détermination non seulement de ce qu'il faut modifier, mais aussi des raisons pour lesquelles une modification est justifiée et du moment où elle doit devenir une orientation persistante. Cependant, les méthodes existantes fondées sur l'expérience peuvent perdre les preuves comportementales et les contextes de tâche qui étayent les modifications. De plus, un résultat de validation global ne fournit qu'un jugement incomplet des modifications qui le composent : des corrections localement étayées peuvent être écartées lorsqu'une révision est rejetée, tandis que des preuves peuvent nécessiter une expérience supplémentaire pour éclairer des mises à jour utiles. À cette fin, nous introduisons EVISKILL, un cadre dirigé par les preuves qui organise les observations d'exécution en cartes de preuves rejouables (« Replayable Evidence Cards ») et synthétise des modifications en les reliant explicitement à leurs contextes de justification. Le rejeu ciblé vérifie ces modifications par une réexécution et fournit un retour en vue de leur correction. Au fil des époques, EVISKILL conserve les preuves et retient provisoirement les modifications étayées en vue d'un affinement ultérieur, tandis que la validation globale régit leur incorporation dans la compétence finale. Des expériences sur trois bancs d'essai interactifs et six modèles LLM de base démontrent l'efficacité de cette approche. Notre code et les détails de mise en œuvre sont disponibles à l'adresse https://github.com/Zhouyaner/eviskill.
One-sentence Summary
Researchers at Jilin University propose EVISKILL, an evidence-driven framework for continual LLM skill evolution that stores execution observations as Replayable Evidence Cards, synthesizes edits explicitly linked to supporting contexts, verifies those edits through targeted replay, and provisionally retains supported edits for further refinement before global validation incorporates them into final skills, with effectiveness demonstrated on three interactive benchmarks across six LLM backbones.
Key Contributions
- EVISKILL is an evidence-driven skill evolution framework that organizes execution observations into Replayable Evidence Cards and links skill edits to their supporting behavioral contexts.
- Replay-guided edit verification tests whether edits produce their intended behavioral effects through re-execution, while cross-epoch evidence propagation retains and refines supported edits even when the overall revision is rejected.
- Experiments on three interactive benchmarks across six LLM backbones show consistent improvements over strong skill evolution baselines, with analyses indicating that evidence-grounded replay filters unsupported revisions and cross-epoch propagation enables prior evidence to support continued refinement.
Introduction
LLM agents increasingly handle interactive tasks such as tool use, web navigation, and long-horizon decision making, where reusable external skills provide procedural guidance without changing model parameters. Prior experience-driven skill evolution methods extract updates directly from execution trajectories, but execution experience is local and context-dependent, so edits can overfit observed cases or capture incidental behavior. Rejected revisions may also contain useful edits that are discarded. The authors introduce EVISKILL, an evidence-driven framework that organizes continual skill evolution into evidence construction, behavioral verification, and cross-epoch refinement. EVISKILL grounds candidate edits in replayable execution evidence, re-executes referenced behaviors to verify intended effects, and preserves uncertain or unresolved evidence for reconsideration across future epochs. Experiments on three interactive benchmarks across multiple LLM backbones show consistent improvements over strong skill evolution baselines.
Method
The authors introduce EVISKILL, a framework that organizes skill evolution around Replayable Evidence Cards. These cards record proposed skill corrections alongside their supporting execution evidence, enabling a structured pipeline for continuous improvement. The overall architecture consists of three interconnected phases: Evidence-Grounded Edit Synthesis, Replay-Guided Edit Verification, and Cross-Epoch Evidence Propagation.
As shown in the figure below:
In the first phase, Evidence-Grounded Edit Synthesis, the system distinguishes between the Validated Skill, which is updated only upon successful global validation, and the Working Skill used for ongoing interactions. At epoch e, the Working Skill SeW combines the latest Validated Skill SeV with a Provisional Edit Ledger Pe, which contains replay-supported edits retained for further refinement:
SeW=SeV⊕Pewhere ⊕ applies the edits sequentially. The Action Agent collects trajectories using SeW, providing the foundational evidence for skill revision. An LLM evidence extractor analyzes the Working Skill and its trajectories to propose corrections, recording each proposal as a Replayable Evidence Card. A card Ci is formally defined as:
Ci=⟨i,Gi,di⟩where i is a persistent identifier, Gi represents supporting trigger ranges from trajectories, and di is the proposed correction. The framework extracts Trajectory Evidence Cards from current interactions and Contrastive Evidence Cards from adjacent-epoch trajectory comparisons to identify persistent deficiencies. These cards are combined into an Evidence Pool, grouped into Evidence Windows based on compatible contexts, and processed by an LLM editor to synthesize a candidate edit set Ue.
The second phase, Replay-Guided Edit Verification, ensures that synthesized edits produce their intended behavioral effects. For each candidate edit u∈Ue, the system retrieves its supporting cards and replays the associated trigger ranges. By reproducing the prefix of the source trajectory, the framework reconstructs the execution state and re-executes the segment under the modified skill SeW⊕u. An LLM evaluator assesses the source and replayed segments, returning a decision to accept, reflect, or reject the edit. If the decision is to reflect, the editor generates a revised edit u′ based on the feedback, which then undergoes another round of replay and evaluation. Successfully verified edits are consolidated into an ordered collection, yielding a candidate revision Se for global validation.
The final phase, Cross-Epoch Evidence Propagation, governs how skills and evidence are updated across epochs. Global validation compares the candidate revision Se against the current Validated Skill SeV on a validation set. If Se demonstrates superior performance, it replaces SeV, the Provisional Edit Ledger is cleared, and the incorporated edits become the new Tracked Edits for future cross-epoch comparisons.
If the candidate revision is globally rejected, the framework employs post-rejection replay to salvage effective edits. Each consolidated edit is replayed independently under the original Validated Skill SeV⊕v to determine if it remains effective without the provisional edits. Edits that pass this isolated evaluation are retained in the updated Provisional Edit Ledger for the next epoch, ensuring that locally beneficial corrections are not discarded due to global incompatibilities. Furthermore, the lifecycle of the Evidence Cards is dynamically managed. Cards supporting accepted edits are archived, while those supporting rejected edits may be marked as stale or protected for future cross-epoch comparisons. After E evolution epochs, the final Validated Skill is frozen for inference, ensuring that only thoroughly validated and replay-supported skills are deployed.
Experiment
The preliminary study identifies two weaknesses in existing skill evolution: edits based on experience often fail to produce intended execution effects, and globally rejected revisions may still contain useful constituent edits. The main experiments evaluate EVISKILL across App-World, ScienceWorld, and ALFWorld with multiple LLM backbones, showing that grounding skill updates in execution evidence yields broad gains and that evolving a skill does not guarantee improvement over its initial version. Ablation and mechanism analyses confirm that replay-guided edit verification improves behavioral corrections, while cross-epoch evidence propagation preserves supported edits and unresolved evidence, with many edits from rejected revisions later incorporated into accepted skills and evidence cards reused across epochs.
Execution-grounded skill evolution yields broad task-success gains across interactive benchmarks, improving over the no-skill baseline in every setting and ranking among the top methods in most comparisons. Several skill-evolution baselines regress on ScienceWorld, falling below both the no-skill and initial-skill accuracy, while the execution-grounded approach avoids this pattern. The results indicate that verifying updates through observed behavior helps preserve useful skills while incorporating new experience. Execution-grounded skill updates achieve the best accuracy in most model-dataset settings and improve over no skill in all reported settings. On ScienceWorld, several evolution baselines underperform the initial skill, with some dropping below the no-skill baseline, whereas execution-grounded updates improve over the initial skill in nearly all settings.
The experiments evaluate execution-grounded skill evolution on interactive benchmarks, testing whether updates grounded in observed behavior improve task success. The approach consistently improves over the no-skill baseline across all settings and ranks among the top methods, while several skill-evolution baselines regress on ScienceWorld and fall below both no-skill and initial-skill accuracy. Execution-grounded updates also achieve the best accuracy in most model-dataset settings and improve over the initial skill in nearly all ScienceWorld settings. These results suggest that grounding skill updates in execution helps retain useful skills while integrating new experience.