The OpenAI Foundation has awarded $40 million to UNC Lineberger Comprehensive Cancer Center to generate data that will help AI models design more effective personalized cancer vaccines. The grant, announced September 15, 2026, aims to solve a core bottleneck: many vaccine candidates fail because the selection of tumor targets remains too arbitrary.
Computational biologist Alex Rubinsteyn and immunologist Benjamin Vincent will lead the effort. They plan to analyze hundreds of de-identified tumor tissue and immune cell samples from three biobanks. The resulting dataset will train machine learning models to identify stronger cancer-cell targets - the proteins on tumor surfaces that flag cancer for immune attack.
"Personalized cancer vaccines are finally starting to show signs of clinical efficacy, but many fail in the development process because they are too arbitrary," said Rubinsteyn, an expert in machine learning and computational medicine at the UNC School of Medicine. "Our goal is to help the world design personalized and more effective cancer vaccines, using the power of rich data and artificial intelligence."
Closing the data gap in cancer vaccine design
The funding comes through the OpenAI Foundation's second science program, Public Data for Health. The initiative backs open-access scientific datasets that researchers and AI tools can use to advance disease research. Jacob Trefethen, Head of Life Sciences and Curing Diseases at the OpenAI Foundation, said the work will produce data at a scale researchers currently lack.
"AI has enormous potential to help researchers design better cancer treatments, but that progress depends on having the right biological data to learn from," Trefethen said. "By giving the broader scientific community a stronger foundation to work from, we hope to open up new possibilities for cancer research and ultimately help more patients benefit from advances that once felt out of reach."
The research tackles two linked problems. First, current methods for picking tumor antigens often miss the most effective targets. Second, vaccine formulations vary widely in their ability to provoke immune responses, and comparative data is scarce. Vincent and Rubinsteyn will run clinical trials comparing multiple vaccine formulations in triple-negative breast cancer patients, measuring how well each raises tumor-specific immune responses.
This work sits at the intersection of AI for Science & Research and AI for Healthcare, two fields where high-quality training data determines what models can accomplish.
How personalized cancer vaccines work
Cancer vaccines are built from a patient's own tumor cells and immune cells. Proteins on the tumor surface act as signals - Vincent compares them to a red flag in a bullfighting ring. A vaccine trains the patient's immune system to recognize those flags and mount a coordinated attack on cancer cells. The challenge is that every patient presents a different set of tumor proteins and immune system characteristics, making target selection a complex data-analysis problem.
"Selecting tumor antigens that are actually good targets, and understanding the relative potency of vaccine formulations for personalized therapy, are big problems in the field which our work will address," Vincent said.
Why this matters for science and research professionals
The UNC project will make its datasets openly available to cancer researchers and oncologists worldwide. For computational biologists and machine learning engineers working in immuno-oncology, this means access to training data that currently does not exist at this scale. The clinical trial component also provides a direct pipeline from computational predictions to biological validation - a feedback loop that remains rare in personalized medicine research. If the approach succeeds, it could establish a repeatable framework for using AI to shorten the time between tumor sequencing and vaccine delivery.
Your membership also unlocks: