Command Palette
Search for a command to run...
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
Chenguang Wang Ming Li Chengrui Fan Jianpeng Chen Han Chen Tianyi Zhou Dawei Zhou
Abstract
AI reviewers can assign diferent judgments to manuscripts that report the same science in diferent wording, potentially rewarding rhetorical optimization over scientific improvement. We formulate Rhetorical Robustness as the joint requirement of stability across content-preserving rewrites and discrimination across papers. We introduce RobustReview, a controlled full-manuscript benchmark with 1,260 manuscript versions, and evaluate 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers diferently. Moreover, the evaluated contentfocused prompting protocol does not consistently improve robustness across backbones. Motivated by these findings, we introduce SciCore, a dual-branch reviewer that averages a full-manuscript judgment with a judgment based on an extracted, structured science core. This design combines manuscriptlevel assessment with a content-normalized view intended to reduce rhetorical sensitivity. In our primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile among the benchmarked reviewers while maintaining competitive human alignment. These results identify rhetorical robustness as a distinct evaluation target and demonstrate the potential of science-core review to improve it.
One-sentence Summary
Researchers from Virginia Tech, University of Maryland, and MBZUAI introduce RobustReview, a controlled 1,260-manuscript benchmark that exposes false robustness, in which rewrite insensitivity coincides with score collapse, and propose SciCore, a dual-branch reviewer that averages full-manuscript and extracted science-core judgments, achieving a leading stability-discrimination profile and competitive human alignment among benchmarked reviewers.
Key Contributions
- The paper formalizes rhetorical robustness as an important requirement for trustworthy AI reviewers, defined as stability across content-preserving rewrites of the same manuscript and discrimination across papers. Human alignment is evaluated as a distinct property.
- RobustReview is a controlled full-manuscript benchmark with 1,260 manuscript versions and 30 reviewer configurations. The benchmark reveals false robustness, where low rewrite sensitivity coincides with score collapse across papers, and shows that human alignment and rhetorical robustness rank reviewers differently.
- The paper introduces SciCore, a dual-branch reviewer that averages a conventional full-manuscript judgment with a content-normalized judgment from an extracted, structured science core. In the primary GPT-5.5 comparison, SciCore achieves a leading joint stability-discrimination profile while maintaining competitive human alignment.
Introduction
Large language models are increasingly used to support scientific peer review, where their judgments can shape what research is accepted and disseminated. Existing evaluations usually focus on whether AI-generated reviews are useful, resemble expert feedback, or match human decisions, but prior work shows that reviewer models can be manipulated by wording changes and meaning-preserving rhetorical revisions. The authors argue that rhetorical robustness is a missing requirement: a reliable AI reviewer should stay stable across rhetorical variants of the same science while still discriminating between different papers. They introduce RobustReview, a controlled full-manuscript benchmark with 60 ICLR 2026 submissions rewritten under 10 rhetorical conditions, and SciCore, a dual-branch framework that combines a conventional full-manuscript review with a content-normalized judgment of an extracted structured record of the paper’s scientific claims.
Dataset
The authors construct RobustReview from 60 anonymized ICLR 2026 submissions with matched arXiv LaTeX sources. The corpus contains 60 original manuscripts and 1,200 rhetorical variants, for 1,260 full manuscripts total.
- Sources and sampling: Eligible papers are stratified by mean human overall-assessment rating. The authors randomly sample 10 papers from each of six score intervals, balancing coverage across human-assessed ratings.
- Variant conditions: They apply 10 rhetorical conditions. Six alter a single dimension: novelty stance, scope framing, evidence framing, contribution salience, technical register, or linguistic complexity. Four complex conditions apply joint rewrites across dimensions, two or three recursive rewrite rounds (R2 and R3), or a reviewer-guided rewrite based on model feedback.
- Generation and scale: Each condition is independently instantiated by GPT-5.5 and Claude Opus 4.8, producing two variants per paper per condition. This yields 60 originals plus 1,200 variants.
- Processing and metadata: All rewrites operate on complete LaTeX projects under content-preservation and structural controls. Automated and human audits indicate that core technical content is largely preserved across five assessed dimensions. Each variant is identifiable by its matched paper family, rhetorical condition, and generator model.
- Usage: The benchmark uses matched paper families so within-family comparisons directly measure rewrite-induced instability. Cross-family comparisons provide the reference for whether within-paper consistency coexists with differentiation among papers. Human judgments on original manuscripts serve as a separate external reference for evaluating reviewer assessments.
- Training splits, mixture ratios, and cropping: Not described in the provided excerpt; RobustReview is presented as an evaluation benchmark.
Method
RobustReview and SciCore
The authors approach rhetorical robustness as a distinct evaluation target. In their formulation, judgments about reported science should remain stable when a manuscript is rewritten to change only its presentation, while the same reviewer should continue to separate manuscripts that report different scientific content. For a paper i, the original manuscript is denoted xi0. Each controlled rewrite setting k specifies both a rhetorical condition and a rewrite producer, generating a variant xik intended to preserve the paper’s reported claims, methods, evidence, results, and conclusions. A reviewer configuration m then assigns a judgment ymik to each realization:
xik=Rewritek(xi0),k=1,…,K, ymik=Reviewm(xik),k=0,…,K.The benchmark requires two complementary properties. Within-paper stability measures whether judgments remain consistent across rhetorical variants of the same paper. Between-paper discrimination measures whether the reviewer still preserves meaningful distinctions across different manuscripts. The joint requirement prevents a trivial constant reviewer from scoring well on stability alone. The authors operationalize this with both direct within-paper metrics and joint stability-discrimination metrics.
To build the benchmark, the authors create matched paper families from anonymized ICLR 2026 submissions with associated LaTeX sources. Eligible papers are stratified by mean human overall-assessment rating, and a balanced sample is drawn across six rating intervals. Ten rhetorical conditions are applied. Six alter one dimension at a time: novelty stance, scope framing, evidence framing, contribution salience, technical register, and linguistic complexity. Four apply more complex rewrites: joint multidimensional rewriting, two or three recursive rewrite rounds, and reviewer-guided rewriting based on model feedback. Each condition is instantiated by two independent rewrite producers, yielding two variants per paper per condition. The rewrite pipeline operates on complete LaTeX projects under content-preservation and structural controls.
SciCore Dual-Branch Review
The SciCore reviewer is designed to reduce rhetorical sensitivity without discarding the full manuscript context. Instead of asking a reviewer to ignore rhetoric while reading the complete paper, SciCore first transforms the manuscript into a content-normalized scientific record. The method uses two complementary branches: a manuscript branch and a science-core branch.
The manuscript branch applies an evidence-focused review protocol directly to the complete manuscript PDF. It preserves the conventional full-paper review context and produces an overall assessment on the ICLR-style scale.
The science-core branch first uses an LLM to extract a structured science core from the manuscript. The extraction focuses on the reported scientific record, including the central research idea and claims, problem formulation, mathematical formulations and derivations, methods and assumptions, experimental or theoretical evidence, reported results, author-stated contributions, reproducibility information, and stated limitations.
The extractor is instructed to write in objective, third-person language and to preserve attribution to the original manuscript. In particular, claims and author-stated contributions must remain attributed to the authors, reported results and numerical evidence must be preserved, and the extractor should avoid subjective assessments, evaluative language, emotional framing, and unsupported interpretations. The extractor does not judge whether a claim is convincing or how the paper should be scored.
The resulting science core is a text-based scientific record, not a shortened rewrite of the original paper. Figures are not retained directly. Tables are transcribed explicitly so that their reported values and comparisons remain available. Equations, mathematical arguments, reported numbers, and other details are retained whenever they can be reliably recovered. If information cannot be located or read reliably, the extractor records that uncertainty rather than reconstructing missing content.
At review time, the science-core branch receives only the extracted record, not the original manuscript. Its protocol adapts the standard review instructions to the structure of the science core. The reviewer is instructed to connect the stated problem and claims with the methods, assumptions, mathematical reasoning, and empirical or theoretical evidence; to assess whether the reported results support the author-stated claims; and to use reproducibility information and stated limitations when judging the strength and scope of the evidence. Information absent from the manuscript can be treated as missing scientific support, while extraction uncertainty is handled separately. The branch uses the same ICLR-style evaluation criteria and scoring scheme and produces written feedback, strengths, weaknesses, questions, and an overall assessment.
The two branches are fused by an unweighted arithmetic mean. Let M(x) be the manuscript-branch overall score and B(x) be the science-core-branch score for a realization x. The final SciCore score is
SSciCore(x)=21M(x)+21B(x).Equal weighting gives a symmetric combination without introducing a tuned fusion parameter. The manuscript branch and the science-core branch remain independently interpretable, allowing the behavior of each component to be evaluated separately from the fused reviewer.
Experiment
The experiments introduce RobustReview, a benchmark that measures rhetorical robustness by matching original manuscripts with content-preserving rhetorical rewrites and evaluating reviewers on within-paper stability, between-paper discrimination, and human alignment. Across 30 existing reviewer configurations, prompting alone does not reliably enforce content-focused review, and low rewrite-induced drift can reflect false robustness when it coincides with weak paper differentiation; human alignment and rhetorical robustness also favor different protocols. The proposed SciCore system combines direct science-core review with manuscript review, and experiments show it improves joint stability and discrimination while maintaining competitive human alignment, with ablations indicating that direct core review, the adapted review policy, and equal-weight fusion contribute to these gains.
SciCore achieves the strongest joint stability-discrimination results among the evaluated reviewer configurations, leading on ICC, SPR, and discriminability while also recording the lowest human mean absolute error. Its human ranking correlation is competitive but not the leading value, and it is not the best on within-paper movement or drift. The benchmark also shows that low within-paper instability alone can reflect output collapse rather than useful reviewing, so stability must be read alongside discrimination. SciCore posts the best ICC, SPR, and discriminability among compared configurations, paired with the lowest human mean absolute error and the second-highest human Spearman correlation. Despite its strong joint profile, SciCore does not lead on MAD or Drift SD, and its human Spearman correlation remains below the strict GPT-5.5 configuration. Near-zero drift in some core configurations is associated with near-zero SPR, indicating within-paper stability can be misleading without paper-level discrimination.
With a shared GPT-5.5 backbone, direct science-core review provides a stronger robustness branch than manuscript review or reconstruction; reconstruction worsens within-paper stability while offering only modest joint-metric gains. Among core-only policies, the adapted protocol best balances stability and discrimination, whereas stricter or persistent policies can reduce variation or robustness. Fusing the adapted core branch with manuscript-strict review improves robustness over manuscript-strict and improves human alignment, ICC, and discriminability over the adapted core alone. Direct science-core review is more robust than manuscript-standard or reconstruction, and reconstruction degrades MAD and drift with limited joint gains. Core-Adapted achieves the lowest MAD and drift and highest ICC and SPR among the core-only configurations, while Core-Persistent worsens robustness relative to Core-Standard. The final fused SciCore configuration improves all five robustness metrics over Manuscript-Strict and improves human alignment, ICC, and discriminability over Core-Adapted alone. Within-paper stability can be misleading without paper-level discrimination, as some core policies show near-zero drift alongside output collapse and near-zero SPR.
Matched science-core embeddings are consistently far more similar than different-paper controls across all extractor models and both rewrite producers. The separation persists across mean, median, and fifth percentile, showing that cores from rewritten versions remain close to their originals while staying distinct from unrelated papers. GPT-5.5 exhibits particularly strong separation, but the trend holds across every tested extractor and producer. Matched similarities remain high across all extractors and producers, while different-paper control similarities drop substantially. The separation is visible even at the fifth percentile, indicating that low-similarity matched cases still exceed most cross-paper comparisons.
Science-core branch behavior is criterion dependent. It often improves soundness stability relative to manuscript-based review, but human alignment does not improve uniformly across criteria. Confidence is a clear failure mode: Core-Standard and Core-Strict assign constant confidence scores, so low drift reflects output collapse rather than meaningful stability, while Core-Adapted restores variation and improves confidence discrimination. Soundness stability tends to improve in the science-core branch compared with manuscript-based review, while soundness human alignment does not improve uniformly. Presentation is marked diagnostic because manuscript-based review and science-core review evaluate different constructs: paper presentation versus clarity and completeness of the extracted scientific record. Core-Standard and Core-Strict assign constant confidence scores and have zero SPR; Core-Persistent shows almost no variation, so their near-zero drift is a false-robustness outcome. Core-Adapted restores paper-level confidence variation and raises confidence SPR, showing within-paper stability cannot be interpreted without discrimination.
The comparison evaluates a manuscript-standard full-paper protocol against a core-adapted configuration that uses the same backbone for science-core extraction and review. Core-adapted review improves ICC for all three backbones, but the broader benefits are backbone-dependent. The evidence supports partial transfer of the science-core branch rather than a universal gain. Core-adapted review increases ICC for all three backbones compared with the manuscript-standard baseline. GPT-5.5 improves the five robustness metrics but reduces human alignment. GPT-5-mini improves MAD, drift SD, and ICC while weakening SPR and discriminability. GLM-5.2 improves all seven reported metrics, and human alignment does not improve consistently across backbones.
The experiments evaluate a science-core review pipeline across multiple reviewer configurations, backbone models, embedding extractors, rewrite producers, and review criteria. The main conclusion is that an adapted and then fused science-core branch yields the most balanced robustness profile, with the strongest overall stability and discrimination and lower human error, while strict or persistent core-only policies can collapse variation and create misleading near-zero drift. Matched science-core embeddings remain highly similar across rewrites and extractors, supporting core consistency, but benefits are criterion- and backbone-dependent: soundness stability improves, human alignment is inconsistent, confidence is a failure mode unless variation is restored, and presentation changes the construct being evaluated. Overall, the results support partial transfer of science-core reviewing rather than a universal gain, and stability must be interpreted together with paper-level discrimination.