Evals replace PRDs as core tool for AI product managers, Anthropic lead says

Anthropic's Dianne Penn says evals, not PRDs, are now the primary tool for AI product managers. She cites a Claude 2.0 fix where 80% of "instruction-following" failures were JSON schema errors, now tracked above 99.9%.

Categorized in: AI News Management
Published on: Aug 23, 2026
Evals replace PRDs as core tool for AI product managers, Anthropic lead says

Anthropic product lead Dianne Penn says the product requirement document is no longer the primary tool for AI product managers. In her view, evals - standardized tests that measure model performance against expected outcomes - have taken its place.

Penn, who leads product for Anthropic's AI Research & Labs team, made the case on Lenny's Podcast. She joined Anthropic in 2023, when the entire product team consisted of five engineers and the API business was run by a single engineer. Since then, she has worked on every model release from Claude 2 to Fable, and helped incubate Claude Code, MCP, Skills, tool use, and reasoning capabilities. Before Anthropic, she worked on Amazon's Alexa AI team and traded high-yield bonds at JPMorgan.

Evals turn vague feedback into actionable signals

Penn's core argument is that in model-driven products, the path to user value has changed. Instead of writing PRDs to describe product vision and align stakeholders, product managers now identify user pain points, understand failure patterns, convert them into reproducible eval sets, and feed results back to researchers for targeted fixes.

"Evals are the new PRDs," Penn said in the interview.

She illustrated the shift with an early case study. During the Claude 2.0 era, users repeatedly said "Claude isn't good at following instructions." After drilling down, the team found that roughly 80% of those failures were actually cases where the model couldn't output JSON in the required schema. They generated 30 to 40 failure cases into an eval set, each containing a prompt, the model's response, and a comparison against the expected golden answer. The eval now runs automatically with every new model version, and the metric has stabilized above 99.9%.

This method requires converting vague user complaints into concrete research signals. When users report that "Claude is hallucinating," Penn said, the team must determine whether it's a tool-use failure, a search or knowledge-integration failure, or an alignment problem. Only at that level of specificity can researchers design targeted evals and measure improvement.

PRDs still have a place

Penn does not argue that PRDs are obsolete. When the problem definition is clear, evals can serve as shorthand. But for ambiguous problems, cross-team alignment, and consensus with legal and safety stakeholders, PRDs remain valuable. For example, before the computer use feature launched, the team lacked a clear set of user pain points, so the product vision section of the PRD helped explore "how to first enable a certain user group to use it effectively."

Penn also stressed that leaders must do hands-on work. Even senior product managers at Anthropic follow the same onboarding plan as early-career employees: understand users, read user feedback with consent, and talk to customers. "Someone who has never personally built an AI product can hardly judge what a great AI product or feature should look like," she said. She keeps ownership of one or two model-related workflows herself to stay current on user needs and model capabilities.

On the emergent nature of model capabilities, Penn warned that the "discontinuous capability jumps" described in scaling laws papers mean a model may suddenly acquire an ability the team hasn't detected. "If you don't have evals, and you don't have the corresponding testing systems, then these capability jumps may have already happened - and you wouldn't even know it," she said.

How Anthropic scaled from 5 to hundreds of product people

Penn identified several turning points in Anthropic's trajectory. The Golden Gate Claude project in early 2024 ran for only 24 hours and reached about 2,000 people, but it showed the team that entirely new user experiences could be shipped at startup speed. The Opus 3.0 release built trust between product and research teams - the company had fewer than 200 people then, and teams collaborated remotely over the holidays to complete training and release. Opus 4.5 mattered because it combined frontier model capability with a frontier product experience; without a vehicle like Claude Code, users couldn't fully perceive the model's intelligence.

Anthropic's release pace has accelerated dramatically. Penn said the company's model releases in the second quarter of this year already surpassed the total of all four model families released in 2024. That pace demands intense collaboration, which she described as a "hive mind" - on the night before a release, even team members not directly involved stay to review blog posts, edit content, and design demos.

On hiring, Penn listed three criteria: first-principles thinking, staying close to details, and a grand long-term vision. First-principles thinking, she said, means "figuring out what should be done to achieve the goal," not "copying what has always been done in the past." Anthropic's most successful researchers tend to be skilled at reasoning through problems and willing to dig into training data and eval results.

Why this matters for managers

For product and engineering managers, Penn's approach offers a concrete template: when model capability itself is no longer scarce, systematically converting capability into user value through a rigorous eval system becomes the key variable determining product success. Managers should ask whether their teams have defined, measurable eval sets for the top user pain points - and whether anyone on the team can personally verify model behavior rather than relying on secondhand reports.

Penn also uses Claude personally in ways managers may find useful. She built a Skill based on the book Crucial Conversations to organize communication strategies before difficult conversations and judge the appropriate level of detail. "It's almost like a personalized coach," she said, arguing that AI's value extends beyond work output to interpersonal effectiveness.

She cautioned that forming one's own opinions before deeply interacting with AI is critical, warning of the risk of degraded independent thinking. "Hard-won judgment" will remain a core human value for the foreseeable future, she said - software engineering has already been transformed by AI, while fields like biology and life sciences are just at the start of that curve.

For managers looking to build these skills, an AI Learning Path for Product Managers covers the shift from traditional product methods to eval-driven workflows. Foundational knowledge of how models behave and fail is also useful; Generative AI and LLM Courses cover the underlying mechanics that make eval design possible.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)