You Already Know How to Do Evals

Editor’s note:

Here is what I think is happening. Traditional research roles are forking into at least three directions. Researchers who build. Researchers who evaluate. And researchers who move further upstream, into the hard-to-scope questions that get answered before anyone builds or evaluates anything. Different work, same underlying skill: deciding what good looks like, then defending that definition to someone who disagrees.​

This one comes from lived experience, and it is a POV I don't see discussed enough.​

I teach one of these. Build Like A Pro runs again Saturday, August 29, and Tuesday, September 1. Three hours, eight seats each. I practice the other two but don't teach them yet, which is part of why I'm writing this.​

TLDR: There is real overlap between traditional UXR work and AI eval work, especially where the definition of "good" is incomplete.


A year ago, I was evaluating a model from Cohere, an enterprise AI company, against competing LLMs across North America, the Middle East, Japan, and Korea. I do not read or speak three of the languages involved.

Almost none of the hard parts were technical.

I was researching safety configuration, in-language model performance, and cultural localization. I started planning in spring 2025 with relatively little modeling knowledge. It was daunting! I go qualitative because the questions I find most interesting are the ones a number can't answer. Someone still has to decide what the number is measuring. I did not review metrics in these languages. The team knew these languages weren't performing as well as they wanted. I procured and directed local partners who were fluent in each market's language and context.

What an eval actually is

Remove the tooling, and an evaluation is one simple sentence.

You take a set of inputs, run them through the model, and score the outputs against a definition of good.

This includes three decisions: Which inputs. What good means. Who decides.

 
Before any eval runs: sampling, criteria, calibration

Before Any Eval Runs: Sampling (Which inputs), Criteria (What good means), Calibration (Who decides). Everything downstream inherits these. None are code. All three are judgement.

 

That is sampling, criteria, and calibration. All three are determined before any evals are run, and those choices inform everything that comes after.

The first decision is the one people skip past. Start there. Which inputs is really a recruiting question. In a study, you would not hand every participant the same generic task. You would find people who actually do this thing, then build a task that has their real work inside it. Templated, but variable. Eval inputs can be built the same way. It takes more effort than grabbing a static set of prompts, and it is the difference between testing the model against what you imagined and testing it against what someone actually does.

The engineering around evals may be new. The underlying judgment work often is not.

The rubric is the instrument

You are probably most familiar with rubrics from school grading systems. A rubric is a definition of good, written down, so that different people applying it (educators) reach the same conclusion (when grading student work).

A rubric is a measurement instrument that answers what good means. And instruments fail in known ways.

They fail on validity when they measure something you did not intend. This is like testing a prototype to learn whether an idea resonates, but measuring participants’ reaction to the visual design instead.​

They fail on reliability when two people apply the same rubric to the same output and disagree. Or when the same person disagrees with herself on Tuesday.

Both failures are invisible in the score. The number or the rating looks clean regardless.

This is where defining “good” got hard. The same output could read as fine, or awkward, or factually wrong, or correct but pitched at the wrong level of formality, or scheduled against prayer times, or blind to a local calendar, or unsafe. It depended on who was reading it, their cultural perspective, and what they were specifically looking for.

Most of that is invisible to a score.

Japanese and Korean both build hierarchy directly into the grammar, so a response can be perfectly accurate and still land as insubordinate or condescending. In Korea, a meeting proposed on a “sandwich day,” the working day “sandwiched” between two holidays, tells the recipient you do not know how the calendar works there. In the Middle East markets I studied, the constraints extend well beyond politics. Religion, gender, sexuality, alcohol, and depictions of the human form all carry regulatory or social weight, and none of it is uniform across the region. We were hearing directly from enterprise LLM buyers and deployers in those markets. A model behaving exactly as designed in one market can create real exposure in another. For buyers, that can determine whether to deploy at all.

The part I like the most

Japanese participants told us the model’s Japanese was polite, well-written, and reliable-sounding. Then the same outputs came back containing Chinese characters and words that do not exist in Japanese. One participant said he could not read it. It was not Japanese.

There are reasons this happens. Japanese uses kanji, characters originally adopted from Chinese, so the two languages share many written characters. A model trained on far more Chinese than Japanese can drift across that shared boundary and produce characters that look plausible but are not readable as Japanese.

Fluent. Polite. Grammatically clean. Unreadable.

This was a model problem. What makes it interesting is that it was also an evaluation problem. A script check would have caught it instantly, but nobody had written one, because nobody knew to look. You cannot write the check (or rubric) until a person tells you what to check for.

The evaluation was missing a dimension that mattered to actual users.

This is the eval work I like most. Not the part a score can surface. The part it cannot.

Why teaching is the training ground

Surprisingly, the best preparation I had for this work was not my design or research chops. It was decades of teaching in accredited institutions.

Specifically, in undergraduate and graduate design programs that assessed student work against rubrics I did not write. Three columns. Does Not Meet. Meets. Exceeds, defined as meeting everything in the middle column and then some. Roughly a dozen criteria underneath them, grouped under headings like critical thinking, conceptual skills, and formal skills.

I did not design that instrument. I applied it. Repeatedly, alongside other faculty, to work that did not always sort itself cleanly into one of the three columns.

That is where you learn what a rubric actually is. Not when you write one, although that experience helps. When you use one somebody else wrote and see where it falls short. Not everything is black and white. Especially in design and research.

Now I write my own rubrics. Here are two criteria from Build Like A Pro, where I teach people to visualize their ideas and build prototypes to “show, not tell”, in that same three-level structure.

 

Real enough to react to

  • Does Not Meet: Viewer responds to the artifact, not the idea. Comments on polish, layout, or fidelity.

  • Meets: Viewer asks a clarifying question about the idea itself.

  • Exceeds: Viewer takes a position. Agrees, disagrees, or names what it missed.

 Answers, sparks, or aligns

  • Does Not Meet: Cannot complete the sentence, or the answer is generic.

  • Meets: Completes it. Designed to answer X, illustrate Y, or test Z.

  • Exceeds: The named question is specific and consequential. Someone's decision depends on it.

Exceeds means meeting everything in the middle level and then some.

 

If I hand those two criteria to two people and ask them to grade prototype builds without me, "real enough to react to" will split them. Deciding whether someone reacted to the artifact or to the idea requires reading intent, and two people will not always interpret it the same way. So the rubric is failing on reliability. The people applying it are not.

The fix is not "find better graders." The fix is to improve the instrument. I add an example under each level and calibrate again. If they still split, I split the criterion in two.

Quick vocabulary.

  • The humans applying the rubric are called raters or annotators. They’re the people who decide how good a response is and whether they decide consistently.​

  • The models doing the same scoring at scale are called judges, or LLM-as-judge.

  • Both are scoring the same outputs against the same definition.

That loop is where the accuracy comes from

This is the same loop a team runs before turning on an automated LLM judge. Humans score first. You check whether they agree. If not, you iterate on the definition until they do. Only then does the machine have a standard to match. It is like resetting a bathroom scale to zero before you weigh anything.

Humans using the model in their own language, with their own context, are often the ones who can tell you the definition itself is wrong. But there is a harder problem hiding inside “what good means”: what if the criterion itself is missing?

There is a particularly dangerous version of this failure: everybody agrees. Sometimes the human raters completely agree, and they are all wrong. Everyone marked the Japanese output as good because nothing in the rubric asked them to check the script. Everyone can apply the rubric correctly, while the rubric itself is missing something important.

Disagreement announces itself. Consensus does not.

Researchers do not just apply definitions of good. They help discover when the definition itself is incomplete.

Here are four common technical terms to learn. Researchers already do all four today.

  • Inter-rater reliability. Two people score the same output. Is their scoring the same? This is like two UXRs independently coding the same session. One codes something as a usability issue; the other calls it a preference. If that keeps happening, the problem may be the coding criteria, not the researchers.

  • Calibration sets. A batch of examples everyone scores together before the real work starts, so disagreements surface early. If you’ve used Tomer Sharon’s Rainbow Spreadsheet, you know the loop: everyone observes against the same set of expected behaviors or observations, then you debrief together, clarify what people meant, and tighten the observations before continuing.

  • Golden datasets. A trusted set where the right answer is already agreed on, used to measure everything else. This is like keeping a set of research examples your team has already discussed and agreed on. You know how these examples should be classified, so you can use them to check whether someone or something applies the criteria the same way.

  • Judge alignment. Does the automated scorer match the humans you already calibrated? This is like checking whether another researcher codes the same material the way your calibrated team would. Except now, the other researcher is a model.

These are different terms, but the underlying judgment work is familiar.

In teaching, this is called grading fairly, and anyone who has done it at scale has run the loop hundreds of times.

This is why researchers, educators, content strategists, and anyone who has ever had to define “good” and then defend that definition to someone who disagreed are further along in this work than they may think.

You already know how to work with a rubric. You already know that defining good is hard. What may be new is the vocabulary, the scale, and the fact that you may be writing it for a market you are not from, about content you do not know.

The market twist

The market is starting to say this publicly.

Not many places yet. But the ones that have are naming it precisely.

  • Expedia is hiring a Principal Quant UX Researcher. The posting asks for experience designing AI evaluation frameworks, including human evaluation protocols and LLM-as-judge validation. It also asks for familiarity with psychometric validity frameworks, naming construct validity, criterion validity, and reliability, applied to measuring AI output quality.

Construct validity. Criterion validity. Reliability. That is psychometrics vocabulary in a UXR job description.

Four is not a large sample size. But it is enough to pay attention, especially when these specific employers are asking for it.

The counter-evidence matters too. At the frontier labs, this work is largely scoped under engineering. Anthropic’s model eval roles have Research Engineer titles, and they call for designing eval methodologies and building high-throughput pipelines in the same job. The titles are confusing because they are. The roles genuinely overlap.

Coming full circle

Recently, I had a conversation with an ML engineer who builds and evaluates multilingual LLM systems in production. I did not know her or her work before. She had just published a piece on making evaluation fast enough to iterate on. Careful work, honest about its own limits. She closed it by saying the real leverage is in well-understood engineering applied with judgment.

We found ourselves approaching many of the same evaluation questions from different directions: how representative inputs are selected, how definitions of good are formed, what happens at the linguistic and cultural edges, and what evaluation systems can fail to capture.

Look at that list again. Which inputs. What good means. The first two decisions from the top of this piece. She got there from the engineering side.

She brought the engineering lens. I brought the research and cultural one.

The conversation reinforced for me how much overlap there can be between research judgment and AI evaluation work.

-Michele​

PS. if you've done eval work, or been asked to, I'm interested in what surprised you. Hit reply. I'm collecting evidence. And if you'd rather build than evaluate, Build Like A Pro runs August 29 and September 1.


Speak up, get involved, and share the love!



Previous
Previous

Tell Everyone: What Swimming Taught a Researcher

Next
Next

Learn the lingo: Usability Test