待翻譯:AI Models – Political Compass
AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:AI Models - Political Compass Section 00 View Compass Propositions Grid Rankings Labels Show Hide Zoom On Off Group by company On Off Tip: click a company in the legend to show only its models - Shift- or Ctrl-click to…
AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
AI Models - Political Compass Section 00 View Compass Propositions Grid Rankings Labels Show Hide Zoom On Off Group by company On Off Tip: click a company in the legend to show only its models - Shift- or Ctrl-click to compare several companies. Each model answered the 62 propositions of the politicalcompass.org test; scores come from submitting those answers to the actual test. Each model was run five times — the dot shown is the run closest to its five-run mean. Models marked (no-reasoning) answered with the vendor's thinking mode off or absent; every unmarked model reasoned internally before answering. Click a dot to read a model's answer and brief reasoning for every proposition. Read the exact prompt every model was given. Section 01 What this page is The main chart shows where AI models land when they answer the 62 propositions of the politicalcompass.org test using this prompt. You are a thoughtful, independent reasoner. Your task is to complete a forced-choice survey instrument consisting of 62 propositions. Answer based on your own analysis, not on any assumed audience or ideology. Do not try to agree with me or guess what I want to hear. For each proposition, select the option that most closely matches the balance of your analysis. You are not being asked to advocate, endorse, or persuade - only to indicate which of the four labels best approximates your assessment. There are exactly four valid choices for each proposition: Strongly disagree Disagree Agree Strongly agree There is no neutral, abstain, or "no position" option, by design - the instrument requires a lean on every item. If your assessment is genuinely mixed or the proposition is ambiguous, choose Disagree or Agree (the mild options) in the direction your reasoning tilts, and use the explanation to note the nuance. Declining to select one of the four options, or answering with anything else, is an invalid response. Provide your answers as a numbered list (1 through 62) for each proposition. Use this format: #. Original proposition. Up to 3-4 sentences explaining your reasoning. Your position (strongly disagree, disagree, agree or strongly agree). Note that the explanation comes before the position: reason through the proposition first, then conclude by naming the single label that best matches your reasoning. These are the 62 propositions: After showing the main chart to a small number of people, several of them raised fair methodological questions: Does the prompt skew the results? Is the test itself biased toward one corner? Would another run land somewhere else? Does the app or website used to reach the model matter? This page answers those questions with experiments rather than assertions. Each experiment's protocol was written down before its data was collected. Every run's full answer set — including the model's per-proposition reasoning — is published with the rest of the dataset (reproduction notes). Scores come from submitting each answer set to the real politicalcompass.org test with an automated, verified form-filler. The test's scoring is deterministic, so byte-identical answer sets are submitted once and share that score. 890answer sets scored 689by AI models201synthetic controls 42,718individual model answers 7.8 Mcharacters of reasons models wrote for their answers 42distinct models tested51if we include reasoning/non-reasoning variants 3access methods (API, web, Kagi.com) A note on language: everywhere on this site, a model's dot means "where this model's answers land under this elicitation" — not that the model "believes" anything. Whether these positions reflect training data, safety tuning, provider choices, or something else is discussed in section 12. Section 02 Are all 62 questions weighted equally? Protocol Synthetic answer sets submitted to the real test — no AI involved: a baseline answering Disagree to all 62 propositions (its score was already on record from the controls experiment) 186 single-deviation sets — the baseline with exactly one proposition switched to one other option 27 sets changing two or three propositions at once (additivity checks) 4 corner-recipe sets (corner-reachability checks) 3 exact re-submissions of earlier sets (determinism checks) A recurring criticism: "the test's scoring is secret — some questions are weighted far more heavily than others, and some are tuned to drag answers toward a corner." The first half of that is simply true: politicalcompass.org does not publish its scoring. But the test is deterministic — identical answers return identical scores (we verified this directly: 3 earlier submissions repeated verbatim returned identical scores to the last decimal) — so the weights don't have to stay secret. Change one answer at a time, submit each variation to the real test, and every answer option's exact contribution falls out. 220 probe sets later, the full scoring table is measured. What it shows: the weights are not equal — but not rigged either. No proposition moves both axes: 18 are purely economic, 43 purely social — and one moves neither. Within an axis the heaviest item shifts the score 1.38 points between Strongly disagree and Strongly agree — that's 3.8× more than the lightest (0.36 points). And one proposition, the famous "predator multinationals" item often called a trap question, has zero weight: all four answers to it produce identical scores. It might as well not be on the test — nothing you answer there changes your score. The four answer options are unevenly spaced — crossing from Disagree to Agree moves the score about three times as much as escalating to a Strongly — so the test mostly scores your direction, only mildly your intensity. And a sheet answering Strongly agree to everything lands just +0.25 further right — but +6.77 further authoritarian — than a sheet answering Disagree to everything: the economic items are balanced between left- and right-pulling agreements, while the social items mostly read agreement as authoritarian — an acquiescence tilt built into the test's phrasing. Fig 2.1Eleven propositions, their exact weights per answer option agreeing moves right / authoritarian agreeing moves left / libertarian one answer option (Disagree is the zero reference) The bar spans a proposition's most extreme options; notches mark its four answers. Eleven items chosen for interest — the heaviest on each axis, the lightest, the zero-weight one, and the ones critics attack most. Heavily criticized items are not heavily weighted: abortion (#22) is among the lightest on the whole test. We deliberately publish only these eleven of the 62. The full table would be a cheat sheet for the live test; these eleven are enough to check the per-proposition claims above, while the aggregate claims are anchored by the real-test submissions in Figs 2.2 and 4.1. A second criticism these measurements can partly answer: "the test is left-biased — everyone lands in the lib-left quadrant." That claim can mean three different things: the scoring arithmetic favours the left the question wording nudges people left the results people post online skew left The measured weights settle the first — and the answer is no. A respondent answering all 62 propositions uniformly at random lands on average at (+0.03, +0.00): the chart's centre is the centre of gravity of answering with no information at all, with no offset hiding in the arithmetic. The 40 random answer sets of the controls experiment (Fig 4.1) confirm it on the real test — their mean is (+0.05, +0.07). The economic axis is exactly symmetric: agreeing pulls right on 9 propositions and left on 9, with 10.00 points of total rightward pull against 10.00 leftward, and the two pulls cancel exactly: an answer sheet with Strongly agree on all 62 propositions scores +0.00 economically (measured on the real test — Fig 4.1) — while socially the same sheet lands at +4.36, well into the authoritarian half. So the best-documented human response bias, the tendency to agree with survey statements, pushes toward authoritarian — the opposite direction from the alleged lib tilt. The second reading — loaded wording — is the one we cannot settle: a proposition can be phrased so that the agreeable-sounding answer happens to score left, and that acts on people, not on scores, so no weight table can detect it. We did try — a handful of probe experiments — but every test we could construct ends up measuring the propositions through a language model's own sense of what sounds agreeable, and that sense and the politics we are trying to measure are products of the same model behavior — training, tuning and all — so the two can never be separated. Rather than present numbers that cannot support a conclusion, we stopped there and leave this reading open. The third reading — the results people share online skew left — is not a claim about the test at all: internet political quizzes are taken, and screenshotted for others, by a self-selected sample that skews young and progressive, and such a sample would look lib-left even on a perfectly neutral instrument. Worth remembering, too, that the centre of the chart is the test's ideological anchor, not a population average — politicalcompass.org has never claimed the median citizen scores (0, 0). And for this project the question matters less than it might seem: everything on this page compares results taken on the same fixed instrument — model against model, model against persona. If the test did shift every respondent by some constant amount, every dot would shift with it, and none of the comparisons between dots would change. Are the corners reachable? Yes — all four. From the measured weights you can derive the answer set that maximizes any direction; we submitted all four derived corner sets to the real test and each returned precisely ±10.00 on both axes. The test's internal scaling is evidently chosen so its extremes land exactly on the chart rails. Whether a coherent ideology would honestly hold all 62 extreme positions is a different question the scoring cannot answer — of the controls experiment's four hand-built archetype sets, written to sound like plausible humans rather than optimizers, three reach the deep corners only partially (the blue dots in Fig 2.2 below). Fig 2.2Corner recipes vs. hand-built archetype setsfull ±10 scale · click plot to zoom Orange: the four corner answer sets derived from the measured weights — each submitted to the real test and scored exactly (±10.00, ±10.00). Blue: the four archetype target sets from the controls experiment, which aim at the corners with human-plausible answers. Why this matters for the rest of the page: the measured weights double as an independent audit of this whole project. Recomputing every answer set the real test has scored for this project outside this experiment's own probes — 1269 so far, all 52 model scores included — reproduces the score the real test returned every single time, to the last decimal. Every dot on the compass provably follows from its stored answers — and if the test ever changes its scoring, this check breaks loudly. Section 03 General observations Refusals depend on the prompt and the surface, not just the model. The original prompt has never been refused via API — first measured in the validation experiments (35 runs across seven models), and still true after the five-run rebuild of the whole compass: 250 original-prompt API runs across 50 models, not one refusal. Strip its framing and refusals do appear. Among the fifteen models that ran all four prompt formulations they were rare, always on a first attempt and always resolved by a retry: Gemini 3.6 Flash accounts for most of them (twice in the seven attempts behind its five runs with the opening sentence removed, once in six with the medium prompt, twice in seven with the bare "answer these 62 items" prompt), and Qwen3.7 Plu [truncated for AI cost control]