Rosenverse
Hands-on AI #2: Understanding evals: LLM as a Judge

Log in or create a free Rosenverse account to watch this video.

Log in Create free account

100s of community videos are available to free members. Conference talks are generally available to Gold members.

Hands-on AI #2: Understanding evals: LLM as a Judge

Wednesday, October 15, 2025 • Rosenfeld Community

This video is featured in the Evals + Claude playlist.

Share the love for this talk
Hands-on AI #2: Understanding evals: LLM as a Judge
Speakers: Peter Van Dijck
Link:

Summary

If you’re a product manager, UX researcher, or any kind of designer involved in creating an AI product or feature, you need to understand evals. And a great way to learn is with a hands-on example. In this second talk in the series, Peter Van Dijck of the helpful intelligence company will show you how to create an eval for an AI product using an LLM as a judge (when we use a Large Language Model to evaluate the output of another Large Language Model). We’ll have a look at how that works, but also dig into why this even works. Are we creating problems for ourselves when we let an LLM judge itself? This talk is hands on; and there will be plenty of time for questions. You will go away understanding when and how to use LLM as a judge, and build some product sense around how the best AI products today are built, and how that can help you use them more effectively yourself.

Key Insights

  • •

    Evals are a foundational feedback loop defining what 'good' means for AI products, helping to measure and improve systems continuously.

  • •

    Evaluating fuzzy, subjective AI outputs requires innovative approaches such as using LLMs as judges to score results.

  • •

    Binary (yes/no) scoring is more reliable than rating scales with ranges because LLMs lack internal memory and consistency.

  • •

    Starting evals early (week one of a project) drastically improves AI product outcomes, but many teams delay due to perceived complexity.

  • •

    High-risk or important tasks should be prioritized for evals instead of attempting broad coverage.

  • •

    Assigning a dedicated owner or 'benevolent dictator' for evals who works closely with domain experts accelerates feedback and quality.

  • •

    Creating a written constitution of principles helps concretize AI behavior goals and guides prompt and model training.

  • •

    Most current eval tooling is too technical, slowing iteration cycles and making expert involvement inefficient.

  • •

    Custom feedback interfaces tailored to expert users significantly speed up evaluating AI outputs in domains like healthcare and law.

  • •

    Diverse perspectives from UX, product, strategy, and domain experts are critical in defining and refining what 'good' means in AI systems.

Notable Quotes

"Evals are everywhere, right? Everybody's talking about evals. It is like one of the key things in developing useful AI products."

"You want to ask an LLM to evaluate the fuzzy stuff because there’s no black and white output."

"LLMs don’t have memory, so rating on a scale from one to five is pretty random. Better to have yes or no answers."

"One of the biggest problems in AI building is evolving your prompts and having a fast feedback loop."

"By starting to categorize risk in detail, you naturally lead to better prompts and better evals."

"A constitution is a very good exercise: write down your system’s principles and values to help guide its behavior."

"Use custom systems for experts to quickly review and rate outputs, making feedback cycles much faster."

"Evals define a shared definition of good with tests to measure it, and that is the secret sauce for building great AI products."

"Model companies are students in a classroom wanting good points—they’re happy to run external expert evals to improve."

"The more I work with evals, the more I think UX and product people need to be involved because of the need for diverse perspectives."

Ask the Rosenbot
Jim Kalbach
Jobs To Be Done
2021 • Enterprise Community
Robin Beers
Navigating organizational systems: Rethinking researcher’s role in driving change
2025 • Advancing Research 2025
Gold
Leah Buley
Closing Plenary: The Crisis of Digital
2020 • Advancing Research 2020
Gold
Catt Small
Moving from Execution to Strategy as a Designer
2022 • Design in Product 2022
Gold
Kelly Dern
AI as a Design Partner: How to Get the Most Out of AI Tools to Scale Your Process
2023 • DesignOps Summit 2023
Gold
Patrizia Bertini
Pushing DesignOps’ Influence into New Global Markets
2022 • DesignOps Summit 2022
Gold
Prayag Narula
HCI 2.0: Humanity Deserves the Attention that UX Research has to Offer
2023 • Advancing Research 2023
Gold
Robin Beers
Beyond Insights: Researchers as Organizational Change Catalysts
2024 • Advancing Research 2024
Gold
Tatyana Mamut
Opening Keynote: Breaking Conway's Law--or How to Work Differently and Not Ship Your Org Chart
2019 • Enterprise Experience 2019
Gold
Louis Rosenfeld
Welcome / Housekeeping
2023 • Enterprise UX 2023
Gold
Frances Yllana
DesignOps Exposed: What do our peers really think of us?
2025 • DesignOps Summit 2025
Gold
John Calhoun
Meters, Miles, and Madness: New Frameworks to Measure the (Elusive) Value of DesignOps
2024 • DesignOps Summit 2024
Gold
Cara Maritz
The Art of Extrapolation
2023 • Advancing Research 2023
Gold
Amy Parness
Scaling Sustainability: Complementary strategies that drive long-term success
2025 • Climate UX Interest Group
Brad Peters
Short Take #1: UX/Product Lessons from Your Industry Peers
2022 • Design in Product 2022
Gold
Dan Willis
Enterprise Storytelling Sessions
2015 • Enterprise UX 2015
Gold

More Videos

Yalenka Mariën

"A big challenge is that a lot of governmental sites are filled with legally correct but totally incomprehensible content."

Yalenka Mariën Marie Mervaillie

Designing for Digital Inclusion in the Belgian Government

December 8, 2021

Megan Nipe

"This veteran-centric approach gives us a foundation to grow outreach and better serve veterans struggling with PTSD over time."

Megan Nipe Lyndsay Booth

Human-Centered Design for Engagement: Maturing from Newsletterville to Personalized, One-to-One Messaging

December 8, 2021

Dr. Jamika D. Burge

"Hearing a participant say thank you for listening is one of the most powerful indicators of meaningful research."

Dr. Jamika D. Burge Mansi Gupta

Advancing the Inclusion of Womxn in Research Practices

September 15, 2022

Jemma Ahmed

"We must switch our mindset from extractive to relational, from creator to also curator."

Jemma Ahmed Robert Fabricant Sean McKay Llewyn Paine Kate Towsey Noah Bond

Theme Panel

March 11, 2025

Louis Rosenfeld

"We've worked with speakers for months to prepare our conference content, so it's expensive to produce, but really good."

Louis Rosenfeld

About the Rosenverse

March 15, 2026

Nalini P. Kotamraju

"I fumed that my idea had been ignored, wondering if it was because I was new, a woman, brown, or had a PhD."

Nalini P. Kotamraju

Two Jobs in One: Being a “Leader who is a Researcher” and a “Researcher who is a Leader"

March 10, 2021

Sarah Barrett

"Visible AI interfaces introduce another place where you can add ambiguity."

Sarah Barrett

AI in Real Life: Using LLMs to Turbocharge Microsoft Learn

February 13, 2025

Jane Davis

"AI is an enabler, not a replacement."

Jane Davis

Strategic Shifts and Innovations in User Research: Navigating Challenges and Opportunities

March 11, 2025

Alba Villamil

"Latino parents framed classrooms as the teacher’s domain and home as theirs, avoiding engagement to respect authority."

Alba Villamil

Stereotyped by Design: Pitfalls in Cross-Cultural User Research

March 30, 2020