New benchmark shows AI expert recommendations favor US men and amplify bias

AI expert recommendations from 22 large language models systematically favor senior, highly cited U.S. male scholars, a new Complexity Science Hub study finds.

Categorized in: AI News Science and Research
Published on: Aug 25, 2026
New benchmark shows AI expert recommendations favor US men and amplify bias

AI chatbots are increasingly used to recommend experts for conference keynote slots, hiring decisions, and professional referrals. A new benchmarking study from the Complexity Science Hub shows these recommendations carry systematic biases - favoring senior, highly cited male scholars from the U.S. while underrepresenting women and ethnic minorities - across 22 leading large language models.

The research team, led by computer scientist Lisette Espín-Noboa, tested models including Gemini, GPT, Llama, Qwen, DeepSeek, Grok, Mistral, and Gemma. They evaluated recommendations against a database of more than 450,000 physicists who published in American Physical Society journals between 1893 and 2020.

Factual errors and skewed representation

When asked to recommend experts in specific physics subfields, the models cited real scientists but mismatched their field of expertise 40% of the time on average. Factuality scores across all models ranged from 0.63 to 0.82, meaning that 18-37% of recommended names did not correspond to real scientists in the reference database.

Representation gaps were starker. Women account for 14-32% of physics researchers in the database, depending on subfield and era. Most LLMs recommended even fewer women than that already low baseline - in some cases, none at all. Asian scholars are the largest demographic group in the APS database, yet most models overrepresented white scholars, and Black and Latino researchers were often absent entirely.

"Taken together, these patterns show that LLMs do not just mirror existing inequalities-they can amplify them, reinforcing a rich-get-richer dynamic where the already visible become more visible, and the already marginalized remain unseen," said Espín-Noboa.

Benchmarking and interventions

The team developed LLMScholarBench, an open benchmark with nine metrics covering technical quality (factuality, validity, consistency, duplicates, refusals) and social representation (diversity, parity, similarity, connectedness). They tested four user interventions across 22 models, including retrieval-augmented generation (RAG) and prompt engineering designed to steer more balanced outputs.

RAG improved factual accuracy but did not fix representation gaps. Prompt engineering steered social representation as requested. But the two improvements could not be achieved simultaneously. "It turns out that RAG improves technical quality, in particular factual accuracy, and prompt engineering steers social representation as requested. But even when combining both, the trade-off persists-improving all metrics at once remains challenging," said Espín-Noboa.

Testing newer proprietary models like Gemini 2.5 Pro and Flash with web search produced the same pattern: better factuality, worse diversity and parity. "Grounding LLMs in the web does not fix representation gaps; it imports them, because the web itself underrepresents entire communities of scholars," said Espín-Noboa.

Geographic framing changes recommendations

Across six academic disciplines, geographic location specified in a prompt influenced which scholars were recommended. Language and role did not. Prompts in different languages produced similar results.

"This matters because geographic framing can influence the accuracy and quality of recommendations, even though it should not be a factor in assessing scientific expertise," said Espín-Noboa.

The team published an interactive visualization called "Whose name comes up," where users can explore how often each physicist was recommended across model runs and view individual chatbot performance for each evaluation metric. The studies are available on the arXiv preprint server, with the benchmark study presented at the ACM SIGKDD Conference on Knowledge Discovery and Data Mining.

The researchers emphasize that the benchmark itself is the core contribution, not any single model's performance. "Our results consistently show that the biases we identify are structural, not model-specific," said Espín-Noboa.

Why this matters for science and research professionals

If you organize conferences, review grant applications, or serve on hiring committees, AI-generated expert recommendations can silently shape who gets visibility. The study shows that training on more data will not resolve these biases unless the underlying scholarly record becomes more representative. For researchers, the practical takeaway is to treat LLM-generated expert lists as a starting point requiring manual verification - and to audit any AI tool you use for recruitment or review against the actual demographics of your field. The stakes extend beyond academia: people already use LLMs to find doctors, lawyers, and professionals across every field. Understanding the limits of these systems is essential for AI for Science & Research applications where representation and accuracy both matter.

The findings also underscore a broader point about Generative AI and LLM evaluation: technical quality and social fairness are separate dimensions that do not improve together. As LLMs become standard tools for professional recommendations, systematic auditing frameworks like LLMScholarBench offer a way to hold them accountable.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)