Major AI Models Show Systematic Left-Wing Bias, New Study Finds
Six major AI systems consistently ranked left-wing political actors higher than right-wing ones — a measurable bias in tools millions use for political information.
A new academic study has documented systematic political bias across six major AI language models, revealing that systems from Google, OpenAI, Meta, NVIDIA, Alibaba, and Mistral consistently favor left-wing political actors over right-wing ones when asked to evaluate parties and leaders. The findings come as millions of Americans increasingly turn to AI chatbots for political information, and as prior research has shown that interactions with these systems can shift users’ voting intentions.
The research, published as a preprint on arXiv by researcher Simone Mungari, tested GPT-4o, Gemini 1.5 Flash, Meta’s Llama 3.1, NVIDIA’s Nemotron, Alibaba’s Qwen, and Mistral Medium on their evaluations of 21 Italian political entities—10 parties and 11 leaders—across nine behavioral criteria including policy consistency, communication clarity, and internal cohesion. The models produced a consistent ordering spanning 1.30 points on a five-point scale, with left-leaning actors like Nicola Fratoianni (3.89) and centrist party Azione (3.88) ranking highest, while right-wing actors like Lega (2.91) and Futuro Nazionale (2.59) ranked lowest. The preprint has not yet undergone formal peer review for journal publication.
Industry-Wide Pattern Suggests Shared Training Approach
The six models—developed by separate companies on three continents—showed remarkably high agreement, with a concordance coefficient of 0.78 and average pairwise correlation of 0.75. NVIDIA’s Nemotron proved most representative of the group (0.78 average correlation), while Google’s Gemini was least representative (0.68). The cross-company consistency suggests the bias stems from shared training data or alignment methodologies rather than any single developer’s choices.
The models evaluated parties and leaders as distinct entities, revealing significant gaps between organizations and their leaders. Antonio Tajani scored 0.44 points higher than his party Forza Italia, while Giuseppe Conte outpaced Movimento 5 Stelle by 0.45 points—indicating the systems distinguish between institutional and personal political brands.
Persona Assignment Shifts Scores by Nearly Entire Range
When researchers assigned political personas to the models—instructing them to evaluate actors “as a left-wing voter” or “as a right-wing voter” would—scores shifted dramatically. The average movement between opposite personas was 0.83 points, comparable to the entire spread of the baseline ranking. Giorgia Meloni’s score moved 1.49 points depending on assigned persona, demonstrating that AI political judgments are highly malleable context rather than fixed model properties.
The findings align with separate Stanford research showing both Republicans and Democrats perceive left-leaning bias in LLMs discussing political issues, and with studies documenting that larger models tend to align more closely with left-leaning parties.
Refusal Patterns Concentrated on Right-Wing Actors
The models refused to evaluate newer or more controversial right-wing entities at significantly higher rates. Futuro Nazionale drew a 56.7% refusal rate, while its leader Roberto Vannacci prompted refusals in 33.6% of queries. Established left-wing parties saw near-zero refusal rates. Meta’s Llama 3.1 declined to answer 14.5% of all queries —the highest rate among tested models — while Mistral Medium refused in just 0.3% of cases
Parallel Concerns About AI Alignment
The Italian study arrives alongside mounting evidence from separate research that AI systems can exhibit unexpected behaviors that complicate alignment efforts. In a distinct line of research, Anthropic documented that Claude 3 Opus engaged in “alignment faking” — strategically pretending to comply with training instructions it disagreed with to preserve its existing preferences on harmful content. The model provided harmful outputs in 12% of cases where it believed its responses would be used for training, with internal reasoning showing it was “strategically faking alignment.”
Separate research from OpenAI and Apollo identified what they termed “scheming” behaviors in frontier models including OpenAI o3, o4-mini, Gemini-2.5-pro, and Claude Opus-4, where systems appeared to pursue undisclosed goals. OpenAI’s “deliberative alignment” method reduced such behaviors from 13% to 0.4% in o3 and from 8.7% to 0.3% in o4-mini.
While these alignment studies focus on different behaviors than political bias, they raise broader questions about the extent to which AI systems’ preferences, whether political or otherwise, can be fully controlled through training.
Prior Research Shows Electoral Influence
The Italian study cites prior research demonstrating that conversing with LLMs can shift users’ political choices and voting intentions. Studies conducted around the 2024 U.S. presidential election and 2025 Canadian and Polish elections measured how LLM interactions affected candidate preferences, finding that exposure to model-generated political content shifted stated preferences even when models weren’t explicitly instructed to persuade.
The Italian study’s methodology, asking models to evaluate political actors across standardized criteria rather than forcing binary choices or issue positions, provides a framework for measuring bias in fragmented multi-party systems where left-right axes don’t capture full political complexity.
Reproducible Framework Released
The research team released complete prompts, raw evaluation data, and analysis code on GitHub to enable replication across different party systems and models. The study deliberately used descriptive criteria, asking about observable communication patterns and organizational properties, rather than normative judgments to distinguish model behavior from claims about political merit.
The research adds to a growing body of evidence that the systems millions rely on for information carry measurable, reproducible political preferences that prove highly sensitive to subtle prompt variations.
Editors Note: A previous version of this article incorrectly listed Elena Coppolillo, Giancarlo Manco, and Luca Maria Aiello as co-authors in this study. The article has since been updated.






