Rosenverse
Hands-on AI #1: Let’s write your first AI eval

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #1: Let’s write your first AI eval

Wednesday, October 8, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #1: Let’s write your first AI eval
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this talk, Peter Van Dijck of the helpful intelligence company will walk you through writing your first eval. You will learn the basic concepts and the tools, and write an eval together. This talk is hands on; you can follow along, and there will be plenty of time for questions. You will go away with an understanding of the basic building blocks of AI evals, and with the confidence that you know how to write one. And more importantly, you’ll build some intuition, some product sense, around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • •

    Evals consist of a task, a golden dataset with known correct outputs, and an evaluator that measures correctness.

  • •

    Manual AI prompt testing is slow and inconsistent; automated evals accelerate and scale evaluation.

  • •

    UX and product teams can and should learn evals as a practical, non-technical skill.

  • •

    Creating your own golden dataset is essential and cannot be outsourced or fully automated.

  • •

    Models are fixed once trained; improvements happen by refining prompts and context design, not retraining the model.

  • •

    Evaluations measure task performance, not the underlying model itself, allowing comparison across models.

  • •

    Outputting a confidence score from models is unreliable due to lack of internal memory and inconsistent scale interpretation.

  • •

    Biases are baked into models during training via evals used in post-training refinement.

  • •

    LLMs can be used to judge other LLM outputs to evaluate tasks with non-binary answers.

  • •

    Effective eval work requires collaboration across data analysts, engineers, subject matter experts, and UX/product teams.

Notable Quotes

"Evals are like a way to define what good looks like."

"The model was baked and once it’s baked, it does not learn again until they bake a new one."

"You need to be looking at the data. Nobody wants to, but that’s core work."

"Without a golden dataset, you have to build the golden dataset yourself."

"We’re not teaching the model anything; we’re improving our prompts and context."

"Confidence scores from the model are not a good idea because the model has no memory."

"Biases are baked in through the evals used during model training and post-training."

"LLMs judging other LLMs might sound crazy, but if you do it right, it works."

"Evals are a product and UX skill; learning them lets you make these systems do what you want."

"There is a large and growing capability overhang in these models we haven’t discovered yet."

Ask the Rosenbot
Dan Hill
Strategic design, slowdown, and the infrastructures of everyday life
2022 • Enterprise Community
Landon Barnes
Are My Research Findings Actually Meaningful?
2022 • Advancing Research 2022
Gold
Lisa Gironda
Opener: Chief of Staff–An unexpected journey
2024 • DesignOps Summit 2020
Gold
Darian Davis
Lessons from a Toxic Work Relationship
2024 • Enterprise Experience 2020
Gold
John Calhoun
Have we Reached Our Peak? Spotting the Next Mountain For DesignOps to Climb
2021 • DesignOps Summit 2021
Gold
Scher Foord
Turn the Ship Around: How to Apply Design Thinking Across Your Organization
2021 • Design at Scale 2021
Gold
Sheryl Cababa
Living in the Clouds: Adopting a Systems Thinking Mindset
2023 • Enterprise UX 2023
Gold
Dianne Que
Real Talk: Proving Value through a Scrappy Playbook
2019 • DesignOps Summit 2019
Gold
Sylvie Abookire
A Civic Designer's Guide to Mindful Conflict Navigation
2022 • Civic Design 2022
Gold
Alana Washington
Theme 1: Introduction and Provocation
2024 • DesignOps Summit 2020
Gold
Abbey Smalley
Today’s Design Ops and Programs Landscape & Career Paths
2023 • DesignOps Summit 2023
Gold
Theresa Neil
Just Build Me a Dashboard!
2019 • Enterprise Community
Greg Petroff
Exit Interview #1: Greg Petroff: From Silicon Valley Executive to Sonoma County Possibilitarian
2025 • Rosenfeld Community
Alicia D. Johnson
Disasters and the 21st Century
2021 • Civic Design 2021
Gold
Bria Alexander
Opening Remarks
2022 • Civic Design 2022
Gold
Maria Giudice
Remaking the Making Company: Moving from Product to Experience
2016 • Enterprise UX 2016
Gold

More Videos

Frances Yllana

"Clear is kind — clarity and kindness in content design create meaningful experiences for all."

Frances Yllana Ann Buechner Jess Jones Betsy Ramaccia

D.E.A.R.R. Diaries (Discipline, Experience, Architecture, Reflection + Revolution)

November 16, 2022

Bria Alexander

"The waters of operations can feel incredibly murky and the path forward may not always be clear."

Bria Alexander

Theme Two Intro

September 8, 2022

Dr. Jamika D. Burge

"Doing no harm is a core tenet of user experience research, whether or not you are trained in IRB processes."

Dr. Jamika D. Burge

Broad Strokes: Connecting Design, Research, and AI to the World Around Us

June 7, 2023

Joshua Graves

"Our brains are wired for efficiency and use heuristics, so they prefer daydreaming over hard conversations."

Joshua Graves

We Need To Talk: Navigating Conversations with Your Boss (Part 1 of 3)

April 14, 2025

Lavrans Løvlie

"Europe has leaned more into public sector service design work, which impacts the kinds of metrics and approaches seen in the book."

Lavrans Løvlie Ben Reason

Ask me anything – Authors of Service Design: From Insight to Implementation

November 19, 2025

Rebecca Buck

"I want researchers in the room in positions of power and influence to help executives see the broader social implications of technology."

Rebecca Buck

Mission: Keep Talent in Research Roles!

March 10, 2021

Sarah Williams

"Creating a set of experience KPIs allows us to measure whether our experiences are simple, honest, and meeting customer expectations."

Sarah Williams

Verizon_A Framework for CX Transformation

January 8, 2024

Bria Alexander

"No need to take notes — our scribe David Nicholson is capturing detailed session notes and insights for you."

Bria Alexander

Opening Remarks

September 9, 2022

Joanna Vodopivec

"Customer obsession is actually one of our key values, which makes my job as a researcher a little bit easier."

Joanna Vodopivec Prabhas Pokharel

One Research Team for All - Influence Without Authority

March 9, 2022