Workday has launched Workday AI Research, a dedicated technical research team focused on building reliable, trustworthy, and efficient AI for enterprise use. The team's findings will shape AI development across Workday's HR, finance, and IT platforms, while also contributing methods and evaluation practices to the broader academic community.
The research group has already produced work accepted at top conferences including the International Conference on Machine Learning, the International Conference on Learning Representations, the ACM Web Conference, and the Association for Computational Linguistics. Their focus areas include persistent agent memory, explainability, multi-agent orchestration, reward overoptimization in AI training, recommendation systems, and adaptive resource control.
"As AI agents evolve to remember context and take action on behalf of employees, enterprises are facing complex challenges around privacy, auditability, efficiency, and enterprise-grade accuracy that off-the-shelf models simply cannot solve," said Gerrit Kazmaier, president of product and technology at Workday. "Workday AI Research is dedicated to solving these exact problems, delivering the rigorous science needed to build intelligent, reliable systems that organizations can actually trust and deploy at scale."
PhD fellowship program
To strengthen ties with academia, Workday is introducing the Workday AI Research PhD Fellowship. Doctoral students working at the intersection of AI and enterprise software can receive $50,000 in annual research funding through an unrestricted gift to their university, plus mentorship from a Workday AI researcher and early access to career opportunities at the company. Details are available at workday.com/ai-research.
Recent findings on AI agents
Workday's researchers have been testing whether AI agents can handle memory, collaboration, and deletion tasks reliably enough for enterprise deployment. Their results point to specific technical improvements - and some cautionary findings.
On memory, the team developed a selective approach that helps AI agents keep useful information while filtering out outdated, duplicate, or unreliable details. In testing, the method delivered 12% higher precision and roughly 8% better overall memory quality, while retaining 97% of the memories that mattered. It also ran about 31% faster than the leading AI-driven alternative.
On multi-agent systems, the researchers found that splitting complex work among specialized agents - one exploring options, another enforcing rules, and a coordinator directing the work - improved accuracy by 5.8%. Every final answer in the study met the defined constraints, suggesting organizations don't have to choose between higher-quality recommendations and strong guardrails.
But the research also surfaced a warning about AI memory. When researchers asked an agent to forget something, a copy of the information was still recoverable from an old summary about one in five times. Fully erasing information required deleting every summary that mentioned it, not just the original record.
These findings connect to broader questions in Generative AI and LLM research about how models store and retrieve information. For professionals tracking developments in AI for Science & Research, the work offers concrete evidence about what current systems can and cannot do.
Why this matters for science and research professionals
For researchers working on enterprise AI, the most relevant takeaway is the forgetting study. The finding that deleted information remains recoverable from older summaries has direct implications for privacy, compliance, and audit work. Anyone building systems that handle sensitive employee or customer data needs to account for the fact that "deleting" a record in an AI agent requires clearing every derived copy - not just the original entry. The memory and multi-agent findings also offer practical evidence on how to structure AI systems for both efficiency and rule-following in production environments.
Your membership also unlocks: