A massive new study found that AI digital twins - models built from a person's past data to predict their choices - systematically misrepresent human behavior. The research, published September 4 in Science Advances, compared 1,784 real people with their AI counterparts across 19 experiments and concluded the technology is not ready for real-world use. For behavioral scientists and researchers who hoped these models could accelerate experiments without risking harm to human participants, the findings are a blunt corrective.
What the study tested
Researchers led by Tianyi Peng of Columbia University created a digital twin for each participant by feeding more than 500 prior answers into a base large language model. The twins were then quizzed on hiring decisions, political attitudes, privacy choices, and news consumption alongside the real people they were meant to mirror.
The results were underwhelming. "The digital twins of today may best be described as funhouse mirrors that systematically distort human behavior," the study authors wrote. The twins performed only marginally better at predicting human choices than the same AI model without any personal data at all.
Five systematic distortions
The paper identified five specific ways the twins got people wrong.
First, the models failed to capture individual variation. Their answers clustered together, making distinct people look nearly identical. Second, the twins leaned heavily on stereotypes, defaulting to broad demographic traits instead of personal characteristics.
Third, an ideological bias emerged: the twins were consistently more optimistic about human behavior and technology than the people they represented. Fourth, the models performed unevenly across populations, producing more accurate results for participants with higher incomes and education levels.
Fifth, the twins were too rational. They displayed greater knowledge than their human counterparts and gravitated toward the logically correct answer, even when the actual person did not.
The technology behind the twins
A digital twin is built by training a large language model on an individual's demographics, personality traits, and past decisions. Proponents have suggested these models could one day negotiate on a person's behalf or make choices for them. Companies see potential for running surveys and polls without recruiting human participants.
This study, however, delivers a sharp reality check. The bottom line from the research team was direct: "Digital twins are not yet 'ready for primetime.'"
Why this matters for science and research professionals
For researchers in behavioral science, psychology, and AI for Science & Research, the findings carry a clear warning. Using digital twins as stand-ins for human subjects in experiments risks baking systematic distortions into results - particularly when studying populations outside high-income, high-education brackets. The technology cannot yet replace the messy, irrational, and varied responses that make human data useful. Until these five distortion patterns are addressed, digital twins remain a research tool to validate against real subjects, not a shortcut around them.
The study appears in Science Advances (DOI: 10.1126/sciadv.aeh8260) and was authored by Tianyi Peng and colleagues at Columbia University.
Your membership also unlocks: