Rosenverse
Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Wednesday, July 23, 2025 • Rosenfeld Community

This video is featured in the AI and UX playlist.

Share the love for this talk
Building impactful AI products for design and product leaders, Part 2: Evals are your moat
Speakers: Peter Van Dijck
Link:

Summary

The secret ingredient for impactful AI products is “evals”—an architecture for ongoing evaluation of quality. Without evals, you don’t know if your output is good. You don’t know when you’re done. Because outputs are non-deterministic, it’s very hard to figure out if you are creating real value for your users, and when something goes wrong, it’s really tricky to figure out why. Simply Put’s Peter van Dijck will demystify evals, and share a simple framework for planning for and building useful evals, from qualitative user research to automated evals using LLMs as a judge.

Key Insights

  • AI product development involves three layers: model capabilities, context management, and user experience, with evals central to experience quality assurance.

  • Automated evals help scale testing of AI with inherently open-ended inputs and outputs, enabling faster iteration cycles with confidence.

  • LLMs can serve as judges (evaluators) of other LLM outputs, which works because classification is cognitively easier than generation.

  • Defining what 'good' means for an AI system is a detailed, evolving process informed by research, domain expertise, and observed risks.

  • A three-option evaluation (e.g., yes/no/maybe) works better than fine-grained scales for consistent automated scoring by LLMs.

  • Synthetic data, generated by LLMs based on manually created examples, efficiently expands dataset breadth and usefulness.

  • Domain experts are essential for tagging data and establishing quality criteria, especially for high-stakes areas like healthcare or legal.

  • Building effective evals requires substantial effort—expect 20-40% of project resources devoted to this work.

  • Cultural differences impact subjective evals like politeness, requiring localization and careful domain definition.

  • AI product quality management is a strategic ongoing commitment, extending beyond initial development into production monitoring and iteration.

Notable Quotes

"AI products almost always have both open-ended inputs and outputs, which makes testing really hard."

"You have to build a detailed definition of what is good for my system to do meaningful automated evals."

"It’s much easier to classify an answer than to generate an answer, and that’s why LLM as a judge works."

"You don’t want to give too many options like rating from one to ten because consistency gets lost between different LLM calls."

"Synthetic data is useful because it’s easier to generate more examples of something you already have than to create entirely new data."

"If you launch in the US and politeness is an issue, first try to fix it with prompts; only if that fails should you build an eval."

"Evals are really your intellectual property—they define what good looks like in your domain."

"Domain experts are crucial for tagging data because users might say ‘that’s great,’ but experts can tell it’s totally wrong."

"You should plan 20 to 40 percent of your project budget on evals—it’s a lot more work than most people expect."

"This is where UX and product strategy bring huge value—defining what good means rather than leaving it to engineers alone."

Ask the Rosenbot
Mariah Hay
BUILD: Discussion
2018 • Enterprise Experience 2018
Gold
Theresa Neil
Just Build Me a Dashboard!
2019 • Enterprise Community
Trisha Terhar
Empathizing with the Empowered: Non-Researcher Responses to Democratization
2022 • Advancing Research 2022
Gold
Roy Opata Olende
How Zapier Uses ‘All Hands Research’ to Increase Exposure to Users
2020 • Advancing Research Community
John Paul de Guzman
10k Screens Later: How We Became a Data-Driven Design Organization
2024 • DesignOps Summit 2024
Gold
Rachael Dietkus, LCSW
AI: Passionate defenses and reasoned critique [Advancing Research Community Workshop Series]
2024 • Advancing Research Community
Dan Willis
Enterprise Storytelling Sessions
2016 • Enterprise UX 2016
Gold
Robert Fabricant
Shifting dynamics: The evolving relationship between researchers, participants, and organizational systems
2025 • Advancing Research 2025
Gold
Sam Proulx
SUS: A System Unusable for Twenty Percent of the Population
2021 • DesignOps Summit 2021
Gold
Peter Merholz
Customer-Centered Design Organizations
2017 • Enterprise Experience 2017
Gold
Steve Baty
Discussion
2016 • Enterprise UX 2016
Gold
Jennifer Fraser
What would Emmy Noether Do? Math, Models and Mulling in UX Research
2023 • Advancing Research 2023
Gold
Kate Koch
Flex Your Super Powers: When a Design Ops Team Scales to Power CX
2021 • DesignOps Summit 2021
Gold
Erin Hoffman-John
This Game is Never Done: Design Leadership Techniques from the Video Game World
2017 • DesignOps Summit 2017
Gold
Sharon Banh
Reimagining research: What does the field need to grow? [Advancing Research Community Workshop Series]
2024 • Advancing Research Community
Ben Davies
Expert Panel: The Principles of Research Repository Design
2022 • Advancing Research 2022
Gold

More Videos

Cennydd Bowles

"You can't just measure ethics like a scientific question; you have to reason differently."

Cennydd Bowles

Exit Interview #2: Rediscovering the ethical heart of design

November 6, 2025

Katie Hansen

"The answers to our most pressing questions aren’t out there waiting to be discovered, but right here hidden in what we already know."

Katie Hansen

Finding the unknown in the known: Harnessing meta-analysis and literature review

March 12, 2025

Sam Proulx

"Everyone already uses accessibility features."

Sam Proulx

Accessibility: An Opportunity to Innovate

March 9, 2022

Alexandra Schmidt

"Designers need better training to work with off-the-shelf enterprise software like Sitecore, Salesforce, and SharePoint."

Alexandra Schmidt

Enterprise UX Playbook

December 1, 2022

Caroline Jarrett

"Linking data quality efforts to AI initiatives can help secure attention and budget for necessary improvements."

Caroline Jarrett

Garbage in, garbage out? Measuring error rates to get ready for AI

January 8, 2026

Meghan Bausone

"Maternal health is embedded within a healthcare industry prioritizing cost control, risk management, and standardization, not maternal safety."

Meghan Bausone

Systems Thinking and Design Innovation: Working with Leverage Points in Rural Maternal Health Systems

April 17, 2026

Shelby Switzer

"It’s okay to be goofy and have fun—creating space for awkwardness and playfulness helps engagement."

Shelby Switzer

Making Space for Community Knowledge-sharing in a Distributed World

December 10, 2021

Bria Alexander

"Being a positive deviant means confronting challenges with limited resources but innovative, community-based approaches."

Bria Alexander Ariel Kennan Charlotte Lee Sarah Brooks Emily Lessard Gordon Ross Joanne Dong

Reflect and Chart Forward

December 10, 2021

Smitha Papolu

"Find who the customer success leader is and get them embedded right in your scrum teams at the beginning, that's huge."

Smitha Papolu Nova Wehman-Brown Melissa Schmidt Adam Menter

Theme 3 Discussion

June 4, 2019