The National Science Foundation has awarded Ming Shao, an associate professor in the Miner School of Computer and Information Sciences, a Faculty Early Career Development (CAREER) grant worth nearly $500,000 to improve how AI models handle visual and written data. The research targets a common frustration: generative AI tools like ChatGPT and Google Gemini often misread images when those images are slightly altered or deliberately manipulated.
Shao's work focuses on making AI systems that process multiple forms of data - text, images, audio, and video - more reliable. His team will study how attackers trick these systems and then build defenses against those tactics.
Why training data fails
AI models learn from massive datasets, but that data is often imperfect. "Data is the driving force behind AI models," Shao said. "But that training data can be contaminated."
Outdated text, grainy photos, or deliberately altered images can all lead to incorrect output. In one example Shao cited, a manipulated image can cause an AI tool to identify a banana as an apple.
Shao recently published a paper in the journal Neural Networks analyzing attacks on vision-language models - systems that process both visual and written data. The paper came out of his NSF-funded research.
"First, we explore different types of attacks, and then we try to make the AI models more robust against such attacks," he said.
Building stronger defenses
Shao plans to train AI models on the types of attacks they may face, making them more resistant to manipulation. He is also testing whether one data form, such as written text, can help a model stay reliable when another form, such as a photo, is corrupted.
His team is developing a continuous learning model that lets AI absorb new data without forgetting what it already knows - a problem known as catastrophic forgetting. "You always need to improve your models," Shao said. "My overall goal is to make AI models robust regardless of the data form."
The work connects directly to Generative AI and LLM development, where reliability issues remain a barrier to wider adoption in research settings.
Shao is also applying the research to the new Applied Artificial Intelligence and Data Science undergraduate program, which he oversees. "I want to extend my knowledge and discoveries from this award to the students," he said. "What we learn from our research can be naturally transferred into the program."
Undergraduate and graduate students are assisting with the research, including two Ph.D. students funded by the grant.
Why this matters for science and research professionals
Researchers increasingly rely on AI to process scientific images, literature, and experimental data. When those tools misread inputs, the cost is more than an inconvenience - it can mean wasted time, flawed analysis, or incorrect conclusions. Shao's work addresses the specific failure modes that make AI for Science & Research unreliable: corrupted training data, adversarial manipulation, and model forgetfulness. For professionals who depend on AI-assisted analysis, the practical takeaway is that defenses against these attacks are being built at the model level, not patched on after deployment.
Your membership also unlocks: