Rosenverse
Hands-on AI #2: Understanding evals: LLM as a Judge

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #2: Understanding evals: LLM as a Judge

Wednesday, October 15, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #2: Understanding evals: LLM as a Judge
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this second talk in the series, Peter Van Dijck of the helpful intelligence company will show you how to create an eval for an AI product using an LLM as a judge (when we use a Large Language Model to evaluate the output of another Large Language Model). We’ll have a look at how that works, but also dig into why this even works. Are we creating problems for ourselves when we let an LLM judge itself? This talk is hands on; and there will be plenty of time for questions. You will go away understanding when and how to use LLM as a judge, and build some product sense around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • Evals are a foundational feedback loop defining what 'good' means for AI products, helping to measure and improve systems continuously.

  • Evaluating fuzzy, subjective AI outputs requires innovative approaches such as using LLMs as judges to score results.

  • Binary (yes/no) scoring is more reliable than rating scales with ranges because LLMs lack internal memory and consistency.

  • Starting evals early (week one of a project) drastically improves AI product outcomes, but many teams delay due to perceived complexity.

  • High-risk or important tasks should be prioritized for evals instead of attempting broad coverage.

  • Assigning a dedicated owner or 'benevolent dictator' for evals who works closely with domain experts accelerates feedback and quality.

  • Creating a written constitution of principles helps concretize AI behavior goals and guides prompt and model training.

  • Most current eval tooling is too technical, slowing iteration cycles and making expert involvement inefficient.

  • Custom feedback interfaces tailored to expert users significantly speed up evaluating AI outputs in domains like healthcare and law.

  • Diverse perspectives from UX, product, strategy, and domain experts are critical in defining and refining what 'good' means in AI systems.

Notable Quotes

"Evals are everywhere, right? Everybody's talking about evals. It is like one of the key things in developing useful AI products."

"You want to ask an LLM to evaluate the fuzzy stuff because there’s no black and white output."

"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."

"One of the biggest problems in AI building is evolving your prompts and having a fast feedback loop."

"By starting to categorize risk in detail, you naturally lead to better prompts and better evals."

"A constitution is a very good exercise: write down your system’s principles and values to help guide its behavior."

"Use custom systems for experts to quickly review and rate outputs, making feedback cycles much faster."

"Evals define a shared definition of good with tests to measure it, and that is the secret sauce for building great AI products."

"Model companies are students in a classroom wanting good points—they’re happy to run external expert evals to improve."

"The more I work with evals, the more I think UX and product people need to be involved because of the need for diverse perspectives."

Ask the Rosenbot
Alla Weinberg
Design Teams Need Psychological Safety: Here’s How to Create It
2022 • DesignOps Summit 2022
Gold
Dr. Jamika D. Burge
The Future of Research: Bridging the Gaps
2021 • Advancing Research Community
James Wieselman Schulman
Research is a team sport: advancing the work when everyone does the research
2026 • Advancing Research 2026
Gold
Jake Burghardt
Stop wasting research: Create new value with insight summaries
2025 • Rosenfeld Community
John Maeda
Making Sense of Enterprise UX
2016 • Enterprise UX 2016
Gold
Kim Fellman Cohen
Measuring the Designer Experience
2019 • DesignOps Summit 2019
Gold
Cara Maritz
The Art of Extrapolation
2023 • Advancing Research 2023
Gold
Sarah Brooks
Theme Three Intro
2022 • Civic Design 2022
Gold
Billy Carlson
Ideation tips for Product Managers
2022 • Design in Product 2022
Gold
Dana Bishop
2022: The Year UX Demonstrates its Business Impact
2022 • Advancing Research 2022
Gold
Anna Poznyakov
Get The Most Out Of Stakeholder Collaboration—and Maximize Your Research Impact
2021 • Advancing Research 2021
Gold
Jorge Arango
Exploding the Notebook: How to Unlock the Power of Linked Notes (2nd of 3 seminars)
2024 • Rosenfeld Community
Andrea Gallagher
The Problem Space
2019 • Advancing Research Community
Abby Covert
Stuck? Diagrams Help
2022 • DesignOps Community
Vicky Teinaki
Short Take #3: UX/Product Lessons from Your Industry Peers
2022 • Design in Product 2022
Gold
Amy Thibodeau
Opening Keynote: Process and Ambiguity
2019 • DesignOps Summit 2019
Gold

More Videos

Dr. Karl Jeffries

"If you value innovation and design, then you're investing in creativity."

Dr. Karl Jeffries

The Science of Creativity for DesignOps

January 8, 2024

Milan Guenther

"When we showed the Iceland Ministry of Foreign Affairs their portfolio, the Foreign Minister said, we have to pivot our innovation investments to early-stage catalytic innovations."

Milan Guenther Benjamin Kumpf

The $212 billion ‘so what?’: unlocking impact in development cooperation

November 20, 2025

Ben Davies

"Research repositories aren’t single sources of truth anymore; data and insights live in hundreds of different apps across companies."

Ben Davies Matt Duignan Andrew Michael Dr. Emily DiLeo

Expert Panel: The Principles of Research Repository Design

March 11, 2022

Phil Gilbert

"We’re showing people how to collaborate like crazy and how to fail gracefully."

Phil Gilbert

A Consistent Culture of Design

May 14, 2015

Prayag Narula

"Building more ethical, responsible, and humanistic forms of technologies requires diverse and interdisciplinary conversations."

Prayag Narula Rida Qadri

HCI 2.0: Humanity Deserves the Attention that UX Research has to Offer

March 28, 2023

Ned Dwyer

"Research ops was actually there before research at TravelPerk, supporting designers first."

Ned Dwyer Emily Stewart James Wallis

The Intersection of Design and ResearchOps

September 24, 2024

Erin Weigel

"We are the shopkeepers of today; it just looks a little bit different."

Erin Weigel

Real-world lessons to improve your conversion rates

June 26, 2024

Samuel Proulx

"You cannot design for the middle; instead, you must create experiences that are customizable and adaptable."

Samuel Proulx

From Standards to Innovation: Why Inclusive Design Wins

November 19, 2025

Cheryl Platz

"Mastery is a core human motivator; people want to understand and feel competent in what they do."

Cheryl Platz

Embrace Your Fun Factor: Game Development Best Practices for Product Design

January 9, 2026