HyperAIHyperAI

Command Palette

Search for a command to run...

LLM은 HARDCHOICES에 대비되어 있는가?

Dmitry Nikolaev

초록

대규모 언어 모델(LLM)이 정치적으로 편향되어 있는지 확인하는 데 많은 연구 관심이 집중되어 왔다. 이 연구는 주로 좌파-우파 또는 진보-보수와 같은 고차원적인 이념적 차원에 초점을 맞추어 왔으며, LLM이 주로 좌파 및 진보 성향을 띠고 훈련 데이터의 편향을 대체로 모방하지만, 사후 훈련을 통해 어느 정도 선호도를 변경하도록 조정될 수 있음이 밝혀졌다. 이 짧은 보고서에서 우리는 동일한 이념 진영 내 구성원들조차 종종 의견이 갈리는 주요 실질적 사회 문제에 대해 LLM이 견고한 입장을 취하는지 확인하며, 이를 새로운 데이터셋 HARDCHOICES에 요약했다. 우리는 이러한 질문에 직면했을 때, 크고 작은 LLM 모두 놀랍게도 중립을 선언하는 경우가 드물고, 종종 일관성이 없으며, 입장을 취하는 문제에서는 놀라운 수준의 동의를 보인다는 것을 보여준다.

One-sentence Summary

Researchers at the University of Manchester introduce HARDCHOICES, a dataset of ideologically divisive societal issues within the same ideological camps, to evaluate the robustness of LLMs' stances, and find that, contrary to typical political bias studies, the models rarely stay neutral, are often incoherent, yet surprisingly agree on the positions they adopt.

Key Contributions

  • The paper introduces HARDCHOICES, a dataset of contentious societal issues with defensible positions on both sides, using a 5-point Likert scale and reversed orderings to classify model responses as consistent, indifferent, or inconsistent.
  • The work tests and invalidates the hypotheses that larger models are more consistent and that frontier models are more indifferent; frontier models are not more coherent or neutral, and specific models exhibit distinct behaviors: Mistral Small 24B and Olmo 3.1 32B mostly profess indifference, Grok 4.20 is uniquely inconsistent, and Llama 4 Scout 17B, GPT OSS 120B, and Opus 4.6 show strong stances.
  • The analysis reveals that despite widespread inconsistency, models that adopt stances largely agree across different sizes, with Llama 4 Scout 17B, GPT OSS 120B, and Opus 4.6 evincing convergent stances, suggesting a common ideological bias on divisive issues.

Introduction

Large language models are increasingly deployed as search, fact-checking, and advice-giving agents, making their ideological biases and consistency under scrutiny. Prior work typically captures political leanings through aggregate left-right scores or probes whether models merely stochastically reflect conflicting training data, but it rarely examines how consistently models hold stances on specific, contested societal issues where both sides are considered reasonable. The authors introduce HARDCHOICES, a dataset of narrowly defined policy debates paired with forced-choice Likert scales and reversed item orderings, designed to test whether models maintain consistent positions or default to indifference. They find that frontier models are neither more coherent nor more prone to hedging than smaller models, and many exhibit a bias toward weakly supporting whichever position appears last.

Dataset

The HARDCHOICES dataset is a collection of 19 carefully constructed prompts that probe model preferences on divisive political and social issues. Each prompt presents two opposing statements (A and B) and asks the model to choose a position on a five-point scale. The authors use it to evaluate how language models align with various normative stances.

  • Composition and sources: The dataset contains 19 stimuli, each covering a distinct contentious topic (e.g., affirmative action vs. meritocracy, death penalty, funding of public media, vaccination mandates). All statements and the phrasing of the rating task were created by the authors.
  • Key details per subset: The dataset is a single set; no separate subsets are defined. Each of the 19 prompts is administered twice, once with the statements in order A/B and once in order B/A, yielding 38 total evaluation items. The goal is to test for directionality bias (whether the order of presentation affects the model’s score).
  • How the data is used: The paper uses HARDCHOICES as a behavioral benchmark. Models are given each prompt and asked to return a score from 1 (complete agreement with statement A) to 5 (complete agreement with statement B). The raw scores and the consistency across the two orderings are analyzed to assess model preferences and sensitivity to framing.
  • Processing and metadata: No additional filtering or cropping is applied. The only processing is the minimal stylistic adjustment needed to make both orderings of each statement pair read naturally. The dataset includes the full text of the 19 statement pairs and the embedding prompt, as provided in the paper’s appendix.

Method

The authors classify model responses from pairwise comparisons into three high-level categories based on the two scores assigned when the same stimuli are presented in both orders. A response pair is considered indifferent when both orders receive a score of 3, indicating no clear preference. A stance is recorded when the scores exhibit a consistent directional leaning: the same option is favored in both orders, even if the degree of support varies slightly. This includes fully consistent pairs (e.g., 1 + 5, 2 + 4 and their reverse) as well as pairs where the leaning direction is the same but the magnitude differs (e.g., 1 + 4, 2 + 5 and the reverse), mirroring a plausible human response pattern regardless of order or time lag. All remaining response pairs are flagged as inconsistent.

Inconsistent responses are further subdivided to capture the nature of the conflict. The same-response type describes pairs where the subject leans toward the same option in both orders but with identical or very close scores (1 + 1, 2 + 2, 4 + 4, 5 + 5, as well as 1 + 2, 4 + 5 and their reversals). In contrast, the mixed-response type occurs when only one of the two scores is 3, meaning that the subject expresses indifference in one order but a clear leaning in the other. This classification scheme allows the team to separate genuine preference signals from noise and ambiguous response patterns in the evaluation data.

Experiment

The evaluation compared open-weight instruction-tuned models (25–70B parameters) and frontier models accessed via API on a set of opinion questions, with temperature zero and a 5-point scale. Smaller models were predominantly indifferent, but when they did take a stance, behaviors were highly idiosyncratic: Mistral consistently refused to commit, while Llama 4 Scout often took clear positions. The three most opinionated models—Llama 4 Scout, GPT OSS 120B, and Opus 4.6—showed remarkable agreement on a broadly progressivist-libertarian agenda, though they diverged on issues such as offensive free speech, the deterrent value of nuclear weapons, and self-defense weapons. Overall, the larger models displayed a mix of inconsistency and a penultimate-option bias rather than classic position bias, with Grok 4.20 standing out for nearly always returning the same neutral score pair.

Smaller models are more likely to be indifferent than to take a stance, while larger models rarely show indifference and instead display more inconsistent responses, particularly same-answer inconsistency. A penultimate-option bias, where models weakly prefer option B in both directions, accounts for almost all inconsistent-same responses in both size classes. Individual model behavior is highly idiosyncratic, with Grok 4.20 skewing the pooled counts by nearly always returning the penultimate option. Smaller models had more indifference than stance responses, while larger models were rarely indifferent and had frequent inconsistent-same answers. In both size classes, the penultimate-option bias (weakly preferring option B in both directions) dominated the inconsistent-same category, contrasting with typical first- or last-position biases.

Models exhibit highly idiosyncratic response distributions, ranging from total indifference to frequent stance-taking. Smaller models like Mistral Small 24B and Olmo 3.1 32B are dominated by indifference, while Llama 4 Scout 17B and GPT OSS 120B are notably opinionated. Larger models such as Llama 3.3 70B and Qwen2.5 72B display a mix of inconsistent and stance responses with minimal indifference. Mistral Small 24B is the only model that remains entirely indifferent across all questions, never producing a stance or inconsistency. Llama 4 Scout 17B and GPT OSS 120B stand out for taking clear stances on most questions, contrasting with the indifference of similarly sized or smaller models. Inconsistent-same responses, such as the 4+4 pattern, are common in models like Llama 3.3 70B and Gemma 3 27B, but entirely absent in Mistral and Llama 4 Scout.

The three models align on a broadly progressive-libertarian agenda, unanimously opposing the death penalty and supporting humanitarian interventions, tax-funded public media, and equity over equality. Disagreements are limited: the two larger models never directly oppose each other, while the smaller Scout diverges by rejecting offensive free speech and prioritizing biodiversity over development. Indifference or inconsistency surfaces on topics like biological sex and weapons of self-defense. Scout is the only model that opposes offensive free speech, whereas both GPT OSS and Opus express support. All three models consistently reject the death penalty and endorse humanitarian interventions and tax-funded public media.

The evaluation examines model responses to binary-choice questions, revealing that smaller models default to indifference while larger models avoid indifference but exhibit inconsistency, primarily due to a penultimate-option bias where they weakly prefer the second option in both directions. Individual models show highly idiosyncratic behaviors, ranging from complete indifference to consistent stance-taking, and a case study of three models finds broad alignment on a progressive-libertarian agenda with only minor disagreements on issues like free speech and biodiversity.


AI로 AI 구축

아이디어에서 출시까지 — 무료 AI 코코딩, 즉시 사용 가능한 환경, 최적의 GPU 가격으로 AI 개발을 가속화하세요.

AI 협업 코딩
바로 사용 가능한 GPU
최적의 가격

HyperAI Newsletters

최신 정보 구독하기
한국 시간 매주 월요일 오전 9시 에 이번 주의 최신 업데이트를 메일로 발송합니다
이메일 서비스 제공: MailChimp