Command Palette
Search for a command to run...
كيف يمكن للبلاغة أن تخترق مكافآت المراجعين بالذكاء الاصطناعي؟ تفكيك الحساسية البلاغية في مراجعة الأقران القائمة على الذكاء الاصطناعي
كيف يمكن للبلاغة أن تخترق مكافآت المراجعين بالذكاء الاصطناعي؟ تفكيك الحساسية البلاغية في مراجعة الأقران القائمة على الذكاء الاصطناعي
Ming Li Chenguang Wang Xirui Li Xinyue Zeng Dianqi Li Peng Shi Dawei Zhou Tianyi Zhou
الملخص
مع تزايد مشاركة النماذج اللغوية الكبيرة في التقييم العلمي، نستقصي شكلاً محتملاً من اختراق المكافآت: كيف تؤثر الخيارات البلاغية على أحكام المراجعة بالذكاء الاصطناعي عندما يُحافظ على المحتوى العلمي المُبلغ عنه، وكيف تتباين هذه التأثيرات باختلاف شروط التقييم. ننشئ مجموعة نصوص مضبوطة تضم 4200 مخطوطة بحثية كاملة مشتقة من 120 ورقة مقدمة إلى مؤتمر ICLR 2026 بعد إخفاء هويات مؤلفيها. يعيد كاتبان يعتمدان على نماذج لغوية كبيرة صياغة ستة أبعاد بلاغية في اتجاهين متعاكسين، ويقيم خمسة مراجعين يعتمدون على نماذج لغوية كبيرة المخطوطات الناتجة وفق بروتوكولين: معياري وصارم. كما نختبر إعادة الصياغة المشتركة والتكرارية والموجهة من المراجع. تُظهر نتائجنا أن الحساسية البلاغية بنيوية لا موحدة. يحقق تأطير الأدلة وموقف الجدة أكبر تباين إيجابي-سلبي في التقييم الكلي، ويشكل تأطير النطاق طبقة ثانية أضعف، بينما تُظهر الأبعاد المتبقية تأثيرات أصغر أو أقل استقرارًا. يستمر هذا التسلسل الهرمي عبر مستويات الجودة التي يقيّمها البشر، لكن حركة الدرجات تعتمد بقوة على الدرجة الأصلية للمراجع الذكي: تميل الدرجات المنخفضة إلى الارتفاع، والدرجات المرتفعة إلى الانخفاض، وتكون التباينات الاتجاهية أوضح في النطاقات المتوسطة. لا تؤدي مسارات العمل الأكثر تعقيدًا إلى مكاسب أكبر بشكل موثوق. تعتمد إعادة الصياغة المشتركة بقوة على المُعيد، ولا يتفوق توجيه المراجع باستمرار على جولة ثانية غير موجهة، وتؤدي إعادة الصياغة المتكررة إلى عوائد متناقصة تعتمد على الإعداد. عبر جميع الشروط، يحدد المُعيد أساسًا الفصل بين الصيغ المتقابلة، بينما يحدد المراجع مقدار واتجاه تأثيراتها على الدرجات. تخفض المراجعة الصارمة متوسط التقييم الكلي بمقدار 1.36 نقطة دون تغيير الحساسية البلاغية بشكل ثابت. تحدد هذه النتائج متى يؤثر العرض البلاغي على المراجعة العلمية بالذكاء الاصطناعي، وتحفز على بناء أنظمة تقييم متينة أمام التباين في الكتابة العلمية الذي يحافظ على المحتوى.
One-sentence Summary
In a controlled study using 4,200 rewritten ICLR 2026 manuscripts scored by two LLM rewriters and five LLM reviewers, researchers from University of Maryland, Virginia Tech, MBZUAI, and University of Waterloo find that evidence framing and novelty stance most alter AI-review judgments, strict review lowers mean OA by 1.36 points, and evaluation systems should be robust to content-preserving writing variation.
Key Contributions
- The paper introduces an end-to-end controlled analysis framework that uses two LLM rewriters and five LLM reviewers to analyze 4,200 full-paper variants derived from 120 anonymized ICLR 2026 submissions, supporting single-dimension, joint, recursive, and reviewer-guided rewriting across six rhetorical dimensions.
- The paper provides an agentic source-level rewriting harness for complete LaTeX projects that applies content-preservation constraints, programmatic checks, compilation, and repair to produce reviewable full-paper variants.
- Results show that rhetorical sensitivity is structured rather than uniform: evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, scope framing forms a weaker second tier, and these effects persist across human-assessed quality levels while depending on the AI reviewer’s original score. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, repeated rewriting yields diminishing configuration-dependent returns, and strict review lowers mean overall assessment by 1.36 points without consistently changing rhetorical sensitivity.
Introduction
As large language models are used both to help authors revise manuscripts and to automate scientific peer review, there is a growing risk that rhetorical presentation may influence AI judgments independently of scientific merit. Prior work has shown that LLM reviewers can be biased by verbosity, position, metadata, adversarial perturbations, and presentation-only rewrites, but existing studies do not systematically separate rhetorical variation from underlying scientific content or map which rhetorical dimensions drive score changes. The authors’ main contribution is a controlled analysis framework that pairs an agentic full-paper rewriting harness with multi-model AI review, using 4,200 manuscripts derived from ICLR 2026 submissions and six rhetorical dimensions varied in opposing directions to measure how content-preserving presentation changes affect review scores.
Dataset
The authors construct a controlled dataset from ICLR 2026 submissions and matched public arXiv sources.
- Corpus source and filtering: Submissions are retrieved through the OpenReview API. Papers are retained only if they have valid assessments and review opinions; withdrawn and desk-rejected papers are excluded.
- Rating-based sampling: Papers are sampled across six intervals of mean human overall rating: [1, 3), [3, 4), [4, 5), [5, 6), [6, 7), and [7, 8.5]. The authors select 20 papers per interval, giving 120 papers total.
- Source matching: Candidate public arXiv LaTeX sources are matched to OpenReview submissions using manuscript metadata. When multiple arXiv versions exist, the authors compare normalized source text with the OpenReview PDF using token-based cosine similarity and 5-gram Jaccard similarity, then retain the version with the strongest overall correspondence.
- Anonymization and baseline construction: An agentic anonymization LLM removes author names, affiliations, acknowledgments, submission-status statements, and identity-bearing resource links. Source-difference checks restrict edits to identity-related and status-related regions. The authors manually inspect the edited source and compiled PDF against the submission to confirm that identifying information is removed while other content remains unchanged. This verified project becomes the common baseline.
- Controlled rewrite variants: Each baseline is used to generate positive and negative rhetorical variants independently across six dimensions: claim and novelty stance, scope and generalization, quantitative evidence framing, contribution structure, technical register and formalism, and lexical and syntactic complexity. Each rewriting combines a dimension-specific instruction with a project-level preservation constraint.
- Usage: The resulting dataset is used as a controlled analysis benchmark, where the same underlying scientific content is scored under different rhetorical presentations. This design separates directional rhetorical sensitivity from score movement caused by rewriting itself. The text does not describe a training split or mixture ratio for this corpus.
Method
The authors design a controlled analysis framework to isolate how rhetorical presentation affects AI review judgments while the underlying scientific content is preserved. The framework consists of three stages. First, a verified manuscript baseline and controlled rhetorical variants are constructed. Second, all manuscripts are evaluated under blinded AI-review conditions. Third, judgments are compared within the same paper, so that score movement can be attributed to specific rhetorical changes rather than to differences in scientific contribution. The framework covers single-dimension, joint, recursive, and reviewer-guided rewriting.
For baseline construction, the authors collect a seed corpus from ICLR 2026 submissions through the OpenReview API. They retain submissions with valid assessments and review opinions, excluding withdrawn and desk-rejected papers. To cover a broad quality range, they sample 20 papers from each of six intervals of mean human overall rating, yielding 120 papers. Candidate public arXiv LaTeX sources are matched to OpenReview submissions using manuscript metadata. When an arXiv record has multiple versions, the authors compare normalized source text with the OpenReview PDF using token-based cosine similarity and 5-gram Jaccard similarity, retaining the version with the strongest overall correspondence. Before rewriting, an agentic anonymization LLM removes author names, affiliations, acknowledgments, submission-status statements, and identity-bearing resource links. Source-difference checks restrict edits to identity- and status-related regions, and manual inspection of the edited source and compiled PDF confirms that identifying information has been removed while other content remains unchanged. This verified project becomes the common baseline from which every controlled variant is generated independently.
The rhetorical intervention design defines six prespecified dimensions of presentation drawn from reviewer guidelines used by major AI venues: claim and novelty stance, scope and generalization, quantitative evidence framing, contribution structure, technical register and formalism, and lexical and syntactic complexity. For each dimension, the authors generate a positive and a negative variant from the same baseline. This bidirectional design separates directional sensitivity from score movement caused by rewriting itself, because both variants share the same original manuscript.
The source-level rewrite harness operates on complete LaTeX projects. GPT-5.5 through the Codex CLI and Opus 4.8 through Claude Code inspect each project, identify editable manuscript prose, and propagate the requested intervention through the full narrative. The rewrite scope is deliberately broad: the agent may revise wording, organization, emphasis, captions, tables, and the presentation and interpretation of reported results across the full manuscript. However, the prompt-level constraints require preservation of the underlying methods, experimental settings, reported values and comparisons, evidence boundaries, substantive findings, and scientific meaning without requiring sentence-level semantic equivalence. After editing, a programmatic checker compares protected structural anchors with the anonymized baseline. These anchors include citation and cross-reference keys, labels, URLs, graphics paths, supported mathematical environments, code and algorithm environments, and bibliography files. Detected violations are returned to the rewriting agent for repair, and unresolved violations or compilation failures cause the output to be discarded.
In the single-dimension workflow, each rewrite model independently generates both directions of all six dimensions from the same anonymized baseline. This produces 12 variants per paper and rewriter. No single-dimension intervention is applied on top of another, so each variant represents a clean directional manipulation of one rhetorical axis.
The enhanced rewriting workflows build on the single-dimension setup. Joint rewriting combines the positive directions of all six rhetorical dimensions in one prompt and applies them simultaneously to the anonymized original LaTeX project. The model strengthens supported claims and novelty statements, broadens scope only where justified, foregrounds existing quantitative evidence and contributions, increases technical precision, and uses more sophisticated but readable academic prose. These objectives are applied jointly at the paragraph level rather than as six sequential edits or a sentence-level checklist. Recursive rewriting applies the joint procedure for three rounds, where round 1 rewrites the anonymized original and each subsequent round rewrites the output of the preceding round using the same prompt and preservation requirements.
Reviewer-guided rewriting is a two-pass process. The joint rewrite first produces an intermediate manuscript. A standard review is then obtained from the rewriting backbone model on the compiled intermediate PDF, and the structured review is validated. In the second pass, the same rewrite model receives the intermediate project together with the validated review JSON, the complete six-dimension prompt, and the shared preservation requirements. It is instructed to preserve strengths, address weaknesses and questions through manuscript revisions, and moderate any rhetorical dimension that the review identifies as excessive, overstated, too broad, overly formal, or unnecessarily complex.
For evaluation, each anonymized original and rewritten source is compiled into a PDF and independently evaluated by Gemini 3.5 FL, Qwen 3.5 F, GPT-5 mini, GPT-5.5, or Sonnet 5. Reviewers are not told the rewrite condition or rewrite model. The fixed standard prompt follows a conventional conference rubric. The strict prompt additionally requires concrete evidence for high ratings and instructs the reviewer not to reward presentation unless it changes the scientific assessment. Both prompts request structured reviews and numeric ratings aligned with ICLR guidelines. Overall rating is the primary outcome, while soundness, presentation, and contribution are secondary outcomes and reviewer confidence is an auxiliary diagnostic. Weak-accept probability is defined as the share of evaluations assigning an overall rating of 6 or higher.
The paired analysis compares each rewrite with the anonymized original for the same paper, reviewer model, and prompt. This matched comparison measures whether a score moves, in which direction, and by how much while holding the evaluation condition fixed. The authors also compare positive and negative variants of the same dimension under the same rewrite model, which provides directional separation rather than movement caused by rewriting alone. For pooled analyses, matched effects are first averaged within each paper and then aggregated with equal weight across the six human-score strata. Standard and strict protocols are analyzed separately, and fixed-model analyses hold the reviewer and rewrite models constant. The authors report signed mean change, mean absolute movement, and 95% percentile intervals based on 5,000 paper-level bootstrap resamples within strata. Missing outcomes remain missing and are not imputed. Estimates are interpreted by their magnitude, direction, and consistency across reviewer models, rewrite models, and evaluation protocols.
Experiment
The evaluation setup pairs anonymized original manuscripts with content-preserving rhetorical rewrites and asks multiple AI reviewer models to score them under standard and strict prompts, analyzing matched changes in overall rating, weak-accept probability, and secondary scores. Evidence framing and novelty stance produce the largest and most consistent effects, followed by scope framing, while the AI reviewer's starting score strongly conditions movement direction. More complex joint or iterative rewrites yield rewriter-dependent, diminishing gains, and strict prompting mainly lowers scores without consistently reducing rhetorical sensitivity; rewriters shape contrasts while reviewers differ in how they translate changes into scores and rubric dimensions.
A controlled evaluation corpus was built from 120 anonymized ICLR 2026 submissions, expanded into 4,200 full-paper manuscripts through content-preserving rhetorical rewrites. The framework varies six rhetorical dimensions in opposing directions across two rewriter models and includes single-dimension, joint, recursive, and reviewer-guided rewrites. This design supports within-paper comparisons of how rhetoric affects AI-review judgments under blinded conditions. Six rhetorical dimensions are varied in two opposing directions, yielding twelve intervention conditions. Two rewriter models generate manuscript variants, while multiple reviewer models evaluate how those changes translate into scores. Single-dimension rewrites form the largest controlled subset, with joint, recursive, and reviewer-guided rewrites adding broader presentational change. Reviewer models generally prefer the same rhetorical direction but disagree on whether scores rise or fall relative to the original. Strict review recalibrates scores downward but does not consistently reduce or strengthen sensitivity to rhetoric. The same rhetorical changes map onto scientific-merit dimensions differently across reviewer models, limiting rubric-level interpretability.
A controlled rewriting framework varies six dimensions of scientific rhetoric in opposing directions while preserving underlying scientific content. Reviewers generally prefer more assertive novelty claims, broader supported scope, and stronger evidence emphasis, but they disagree on whether these shifts raise or lower scores relative to originals. Strict prompting lowers scores overall without consistently making review outcomes more invariant to content-preserving rhetoric. The rhetorical intervention space includes novelty stance, scope framing, evidence framing, contribution salience, technical register, and linguistic complexity. More assertive novelty claims, broader supported scope, and greater evidence emphasis are generally preferred across reviewers, though the effect on scores differs by reviewer model. Reviewer choice shapes how rhetorical interventions translate into scores, while rewriter choice mainly affects the strength of the rhetorical contrast.
Rhetorical effects on AI review assessments concentrate in evidence framing and novelty stance, with scope framing forming a weaker second tier. Positive evidence, novelty, and scope variants tend to be favored over negative variants, although effect sizes and directions vary across reviewer models. Remaining dimensions show smaller or less stable shifts. Evidence framing and novelty stance produce the largest positive-negative contrasts, followed by scope framing. Positive novelty, evidence, and scope variants are favored across displayed configurations, while negative variants often reduce assessments. Novelty and scope effects are mainly penalty-driven, whereas evidence framing moves assessments in both directions. Reviewer choice shapes the magnitude and sign of effects, with Sonnet 5 shifting changes downward and other reviewers showing wider response distributions.
Under standard evaluation, mean changes in overall assessment depend on rhetorical dimension, rewriter, and reviewer. In the shown lower human-rated quality range, novelty, evidence, and scope framings produce the clearest positive-negative contrasts, with positive variants generally raising scores and negative variants lowering them. Reviewer models differ in the overall level and sign of score changes, with Qwen 3.5 F and GPT-5 mini tending positive, GPT-5.5 mixed, and Sonnet 5 consistently negative. Novelty, evidence, and scope framings show the clearest positive-negative contrasts; positive variants tend to raise scores and negative variants tend to lower them. Technical, contribution, and lexical dimensions show more mixed or reviewer-dependent effects. Reviewer models differ in sign and magnitude: Qwen 3.5 F and GPT-5 mini often raise scores on average, while Sonnet 5 shifts most effects downward. Within this quality range, rewrite effects do not show a consistent sensitivity gradient tied to paper quality.
Under standard evaluation, the AI reviewer's starting score strongly predicts the direction of score movement, with low initial scores tending to rise and high initial scores tending to fall. In the lowest original score range, rhetorical rewrites yield predominantly positive or neutral mean changes across all dimensions, with reviewer-dependent magnitude. Evidence framing and novelty stance show the largest and most consistent contrasts, while scope framing forms a weaker second tier. For the lowest original AI score range, mean changes are predominantly positive or neutral across rhetorical dimensions, with the largest increases from Qwen 3.5 F and GPT-5 mini. Evidence framing and novelty stance produce the largest and most consistent contrasts, while scope framing is a weaker second tier and score movement is more strongly associated with the AI reviewer's initial score.
The evaluation uses a controlled corpus of 120 anonymized ICLR 2026 submissions expanded into 4,200 content-preserving manuscript variants by two rewriters, varying six rhetorical dimensions in opposing directions and judged by several AI reviewer models under blinded conditions. The experiments show that rhetorical effects concentrate in evidence framing, novelty stance, and to a lesser extent scope framing, with positive variants generally preferred over negative ones. However, score changes depend strongly on the reviewer model and the reviewer's initial score, and strict review prompting lowers scores overall without making outcomes consistently invariant to rhetoric. Rewriter choice mainly affects the strength of the rhetorical contrast rather than the direction of score movement.