WorkshopDoctorDocs

EVALUATE

Measure Learning, Not Happiness

Jason Wright, PhD

A workshop can score 9 out of 10 on satisfaction and produce almost no behavior change a month later. That isn't a participant-motivation problem. It's a measurement problem: the workshop measured how the session felt, not whether it produced what it was supposed to produce.

EVALUATE is the phase that produces the evidence — the fifth of five phases in the Workshop Doctor Roadmap. Its job is to replace satisfaction measurement with evidence measurement: signals gathered Before the session, During it, and After it about whether the transformation actually happened.

When EVALUATE runs as evidence-gathering, the next workshop builds on what actually worked. When EVALUATE runs as a satisfaction survey, the next workshop gets iterated on the basis of feelings — and implementation stays exactly where it was.

Why EVALUATE Matters

Measurement is what closes the loop between building a workshop and building a better one. Without it, each new cohort runs on the creator's intuition about what worked last time. With it, each new cohort runs on evidence — what specifically landed, what specifically didn't, and what specifically needs to change.

The mistake most workshops make is choosing a measurement that feels meaningful but doesn't actually tell you whether the workshop changed anything. Satisfaction scores feel meaningful. They produce numbers, they trend over time, and they make stakeholders feel informed. What they measure is whether the session was a pleasant experience — not whether the participant can now do something they couldn't do before.

Evidence measurement is different. It looks at what the participant produced — the draft positioning statement they wrote, the pricing framework they defended, the work sample they turned in. It looks at behavior a month later: are they doing the thing the workshop was supposed to enable? These are harder signals to collect than a survey rating, and they actually tell you whether the workshop worked.

EVALUATE is the phase that makes the Roadmap self-correcting. Each workshop produces evidence about itself; the next workshop gets built against that evidence. Without EVALUATE, the loop doesn't close — and every cohort runs the same session under the same assumptions.

Evidence Throughout, Not Ratings at the End

Evidence measurement asks one question: after the session, can the participant do something they couldn't do before? Every EVALUATE decision serves that question. Feedback forms, comprehension checks, work-sample reviews, behavior follow-ups — all of it is instrumented against behavior change, not against session feel.

The specific evidence takes different shapes at different moments. Before the session: baseline readings on what the participant can already do, so post-session change is legible. During the session: the comprehension checks DEVELOP built and DELIVER runs — real-time signal on whether concepts are landing. After the session: follow-ups that measure behavior in the participant's actual work a week, a month, and a quarter later.

Satisfaction scores aren't useless. They're a secondary signal — downstream of whether the session produced behavior change, not a substitute for it. A workshop with high satisfaction and high evidence-of-change is a good workshop. A workshop with high satisfaction and low evidence-of-change is a pleasant event. The rating doesn't tell you which one you ran.

EVALUATE is presented as the fifth phase for clarity, but the work of evaluation happens throughout all five. PLAN sets outcomes that evaluation can actually measure; DESIGN builds activities that produce observable evidence; DEVELOP builds the comprehension checks and debrief structures EVALUATE reads; DELIVER runs those checks in real time. The dedicated EVALUATE phase is where the signal gets synthesized and turned into iteration — but the signals are collected throughout.

You are a master of your content. The EVALUATE-phase work is a different craft — the discipline of reading the session against evidence of what changed, and letting that evidence shape the next build. That discipline is what keeps the Roadmap honest across cohorts.

The Core Questions EVALUATE Answers

Do you know what changed? EVALUATE measures behavior change — not engagement, not satisfaction, not how the room felt. The question is specific: a week after the session, a month after, a quarter after, is the participant doing something in their work that they couldn't do before? If the answer isn't a specific observable behavior, EVALUATE hasn't measured what it was supposed to measure.

Are you gathering signal Before, During, and After the session? Evaluation that runs only at the end produces late, shallow data. The methodology treats evaluation as a three-moment discipline: Before the session for baseline readings on what the participant can already do; During the session for real-time signal from the checks DEVELOP built and DELIVER runs; After the session for follow-ups that measure behavior in the participant's actual work. Each moment produces a different kind of signal, and all three are needed to see what the workshop actually changed.

Are you reading what participants produced, or what they reported? Work samples are evidence. Self-reports are opinions. A participant who writes a draft positioning statement produces an artifact the coach can read directly; a participant who rates their confidence on a 1-to-5 scale produces a number that correlates weakly with actual capability. EVALUATE's strongest evidence comes from the artifacts the session produced — the templates filled in, the frameworks applied, the decisions made.

Is the feedback form doing measurement work, or satisfaction work? A short feedback form can do real measurement — five to seven questions focused on engagement, clarity, and the specific outcomes the session was built for. A satisfaction survey of the same length asks different things: "how would you rate this experience," "was the facilitator engaging," "would you recommend this to others." Same format, completely different data. The measurement form asks about what the participant can now do; the satisfaction form asks about how the participant felt.

When changes happen in the next build, are they driven by evidence? Iteration without evidence is guesswork. A creator who adjusts the next cohort based on one loud participant's feedback, or on a gut feeling about what didn't land, is iterating on noise. Evidence-driven iteration looks at patterns — in comprehension-check responses, in work-sample quality, in behavior follow-ups — and changes the design against those patterns. Evidence narrows what to change; emotion widens it.

How to Recognize When EVALUATE Was Reduced to Ratings

Observable signs that EVALUATE is measuring satisfaction instead of evidence:

  • The only post-session data is a rating form with 1-to-5 scales on experience, facilitator, and likelihood to recommend
  • The feedback form has ten or more questions, and most of them ask about how the session felt
  • There's no before-session baseline, so post-session behavior has nothing to compare against
  • The comprehension checks DEVELOP built aren't being read after the session — the data gets collected and forgotten
  • "Did you like it?" is tracked across cohorts while "did you do something with it?" isn't
  • No follow-ups happen a week, month, or quarter after the session to see what the participant is actually doing in their own work
  • Iteration decisions between cohorts get made on the basis of one or two loud comments rather than patterns across the room
  • Changes to the next build are justified by "it felt like the energy dropped there" rather than by specific data from a comprehension check or work sample
  • Satisfaction scores are the success metric stakeholders see; behavior change isn't reported on
  • Feedback gets collected but doesn't feed back into the next cohort's design decisions

These aren't measurement-craft problems. They're evidence that EVALUATE is running as a post-session ritual instead of as the feedback loop that sharpens the next build.

How to Work the EVALUATE Phase

Decide the measurement points at DESIGN time, not after the session. The Before baseline, the During checks, and the After follow-ups all get planned when the session is being designed. Asking "how will I know this worked?" during DESIGN is what produces evaluation that runs as part of the session, not evaluation bolted on as a survey at the end.

Build the measurement instruments into DEVELOP's assets. The Before baseline can be a pre-work task that doubles as a diagnostic. The During checks are the comprehension prompts DEVELOP already builds. The After follow-up is a short structured prompt participants can respond to a week later. Measurement that's built into the package is measurement that actually runs.

Read what participants produced. The template they filled in. The decision they made. The draft they wrote. The work sample they turned in. These are denser evidence than any rating scale. EVALUATE's real reading work happens on these artifacts — the survey is a supplementary signal, not the primary one.

Follow up on behavior. A week after the session: is the participant doing the thing the workshop was supposed to enable? A month after: is it sticking? A quarter after: is it producing the outcomes the participant signed up for? These questions are harder to answer than "how would you rate the workshop," and the answers tell you what the workshop actually changed.

Design a short, focused feedback form — five to seven questions on engagement, clarity, and the specific outcomes the session was built for. Long forms produce low completion rates and watered-down data. Short forms focused on the outcomes the workshop committed to produce data the creator can act on. Satisfaction questions can sit on the form, but they shouldn't be the headline — what the participant can now do is the headline.

Iterate on patterns, not on loudest voices. Individual comments are one data point. Patterns across multiple participants — in comprehension-check responses, in work-sample quality, in behavior follow-ups — are what justify changes to the next build. The discipline is to look at signal density, not signal volume.

None of this requires a research budget or a dedicated evaluation team. It requires the discipline to plan the measurement at DESIGN time, build it into DEVELOP's assets, run it during DELIVER, and read it honestly after the session.

Where this goes next: EVALUATE's evidence becomes the input for the next build. A workshop that measured what actually changed has a starting point the next cohort's PLAN can sharpen against. See Mark's Launch Lab, Before and After for the full arc of one workshop walked through all five phases — or return to PLAN to start a sharper build.

Read this next:

  • Mark's Launch Lab, Before and After — the full arc of one workshop walked through all five phases; the before/after transformation that the phase pages describe at phase level

Foundations of this concept:

Where this leads:

  • EVALUATE closes the loop back to PLAN. What EVALUATE measures becomes the specificity the next PLAN sharpens against. The methodology is self-correcting only when EVALUATE actually runs.

Frequently Asked Questions

Why do most workshop evaluations tell you almost nothing useful?

Most workshops evaluate the wrong thing. A satisfaction survey at the end of the session asks how the experience felt — and produces numbers that correlate weakly with whether the participant can now do something they couldn't do before. The signals that actually answer the workshop's job sit elsewhere: baseline readings before the session, comprehension-check responses and work samples during it, and behavior follow-ups a week, a month, and a quarter after. Iteration on satisfaction data adjusts how the room feels; iteration on evidence data adjusts what the room produces. The two look similar on a survey form and do completely different work for the next cohort.

What is formative vs. summative assessment in a workshop context?

Formative assessment runs during the session: comprehension checks at transitions, breakout deliverables read in real time, structured share-outs that surface what's landing and what isn't. Summative assessment runs at the end: work samples, audit rubrics, follow-up behavior measurement that compares before and after. The Workshop Doctor methodology runs both as part of one Before / During / After discipline. Baseline readings before the session set what change is being measured against. In-session signals tell the workshop creator whether to pause or move on. After-session follow-ups tell whether the change actually translated into behavior in the participant's own work. Satisfaction surveys at the end aren't summative assessment — they're vibe checks, and they shouldn't be the headline.

How do I know if my workshop actually worked?

One question answers it: a week after the session, a month after, a quarter after, is the participant doing something in their work that they couldn't do before? If you can name a specific observable behavior — they're using the pricing framework in actual sales calls, they're shipping the email sequence they drafted, they're closing the conversion conversation differently — the workshop produced what it set out to produce. If you can't name one, satisfaction scores aren't going to tell you. Build the follow-up into the program from the start: structured check-ins keyed to the specific artifacts the session produced, not generic "how's it going?" outreach. The artifacts are the evidence.

Tagged

Free resource

See where your workshops stand

The Workshop Health Check pinpoints exactly where your workshops are losing impact — and gives you a clear path forward.

Take the Health Check →