Rosenverse
Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Building impactful AI products for design and product leaders, Part 2: Evals are your moat

Wednesday, July 23, 2025 • Rosenfeld Community

This video is featured in the AI and UX playlist.

Share the love for this talk
Building impactful AI products for design and product leaders, Part 2: Evals are your moat
Speakers: Peter Van Dijck
Link:

Summary

The secret ingredient for impactful AI products is “evals”—an architecture for ongoing evaluation of quality. Without evals, you don’t know if your output is good. You don’t know when you’re done. Because outputs are non-deterministic, it’s very hard to figure out if you are creating real value for your users, and when something goes wrong, it’s really tricky to figure out why. Simply Put’s Peter van Dijck will demystify evals, and share a simple framework for planning for and building useful evals, from qualitative user research to automated evals using LLMs as a judge.

Key Insights

  • •

    AI product development involves three layers: model capabilities, context management, and user experience, with evals central to experience quality assurance.

  • •

    Automated evals help scale testing of AI with inherently open-ended inputs and outputs, enabling faster iteration cycles with confidence.

  • •

    LLMs can serve as judges (evaluators) of other LLM outputs, which works because classification is cognitively easier than generation.

  • •

    Defining what 'good' means for an AI system is a detailed, evolving process informed by research, domain expertise, and observed risks.

  • •

    A three-option evaluation (e.g., yes/no/maybe) works better than fine-grained scales for consistent automated scoring by LLMs.

  • •

    Synthetic data, generated by LLMs based on manually created examples, efficiently expands dataset breadth and usefulness.

  • •

    Domain experts are essential for tagging data and establishing quality criteria, especially for high-stakes areas like healthcare or legal.

  • •

    Building effective evals requires substantial effort—expect 20-40% of project resources devoted to this work.

  • •

    Cultural differences impact subjective evals like politeness, requiring localization and careful domain definition.

  • •

    AI product quality management is a strategic ongoing commitment, extending beyond initial development into production monitoring and iteration.

Notable Quotes

"AI products almost always have both open-ended inputs and outputs, which makes testing really hard."

"You have to build a detailed definition of what is good for my system to do meaningful automated evals."

"It’s much easier to classify an answer than to generate an answer, and that’s why LLM as a judge works."

"You don’t want to give too many options like rating from one to ten because consistency gets lost between different LLM calls."

"Synthetic data is useful because it’s easier to generate more examples of something you already have than to create entirely new data."

"If you launch in the US and politeness is an issue, first try to fix it with prompts; only if that fails should you build an eval."

"Evals are really your intellectual property—they define what good looks like in your domain."

"Domain experts are crucial for tagging data because users might say ‘that’s great,’ but experts can tell it’s totally wrong."

"You should plan 20 to 40 percent of your project budget on evals—it’s a lot more work than most people expect."

"This is where UX and product strategy bring huge value—defining what good means rather than leaving it to engineers alone."

Ask the Rosenbot
Jennifer Strickland
Fireside Chat: How Design Addresses a World on Fire
2022 • Civic Design Community
Carla Casariego
DesignOps in Wonderland
2019 • DesignOps Summit 2019
Gold
Josh Clark
Sentient Design: New Postures for AI-Mediated Experiences (2nd of 3 seminars)
2025 • Rosenfeld Community
Aurobinda Pradhan
Introduction to Collaborative DesignOps using Cubyts
2022 • DesignOps Summit 2022
Gold
Kit Unger
Theme 3 Intro
2022 • Design at Scale 2022
Gold
Emily Danielson
“I mean, I can lift a shovel”: Design Skills in Disaster Response
2022 • Design at Scale 2022
Gold
Ned Dwyer
Right horses for the right courses – how and when to democratize research
2025 • Advancing Service Design 2025
Gold
Milan Guenther
A Shared Language for Co-Creating Ambitious Endeavours
2023 • Enterprise UX 2023
Gold
James Rampton
The Basics of Automotive UX & Why Phones Are a Part of That Future
2024 • Rosenfeld Community
Sean McKay
Coexisting with non-researchers: Practical strategies for a democratized research future
2025 • Advancing Research 2025
Gold
Hana Nagel
Turning Research Ripples into Waves
2018 • DesignOps Summit 2018
Gold
Craig Brookes
"Just Make it Look Good" and Other Ways We're Misunderstood
2021 • Design at Scale 2021
Gold
Chris Hammond
Embedding sustainability into enterprise design and development: A journey towards "sustainability consciousness"
2025 • Climate UX Interest Group
Amahra Spence
Designing for Liberation, Rehearsing Freedom
2022 • Civic Design 2022
Gold
Anna Avrekh
Diversity In and For Design: Building Conscious Diversity in Design and Research
2021 • Design at Scale 2021
Gold
Adam Cutler
People + Places + Practices = Outcomes
2016 • Enterprise UX 2016
Gold

More Videos

Frances Yllana

"Government might expand its self-perception from service provider to facilitator who partners with constituents."

Frances Yllana Ann Buechner Jess Jones Betsy Ramaccia

D.E.A.R.R. Diaries (Discipline, Experience, Architecture, Reflection + Revolution)

November 16, 2022

Bria Alexander

"You’ll hear tactile guidance on defining ROI."

Bria Alexander

Theme Two Intro

September 8, 2022

Dr. Jamika D. Burge

"Only 12% of AI researchers globally are women, and 6% of professional software developers are women of color."

Dr. Jamika D. Burge

Broad Strokes: Connecting Design, Research, and AI to the World Around Us

June 7, 2023

Joshua Graves

"Anger doesn’t mean you have to yell—it’s about speaking clearly to the problem you see."

Joshua Graves

We Need To Talk: Navigating Conversations with Your Boss (Part 1 of 3)

April 14, 2025

Lavrans Løvlie

"Many new adopters use service design artifacts as a tick-box exercise rather than embracing the mindset change required."

Lavrans Løvlie Ben Reason

Ask me anything – Authors of Service Design: From Insight to Implementation

November 19, 2025

Rebecca Buck

"If your recreational reading is about prison camp survivorship or hostage negotiations, you might have burnout."

Rebecca Buck

Mission: Keep Talent in Research Roles!

March 10, 2021

Sarah Williams

"We want these principles to be broad and channel-agnostic so they can be applied across all teams, not just designers."

Sarah Williams

Verizon_A Framework for CX Transformation

January 8, 2024

Bria Alexander

"We could not have this conference without you all — the curation team relies heavily on audience participation."

Bria Alexander

Opening Remarks

September 9, 2022

Joanna Vodopivec

"Slack channels with live feeds and tagging let busy developers catch key observations asynchronously."

Joanna Vodopivec Prabhas Pokharel

One Research Team for All - Influence Without Authority

March 9, 2022