WorkshopDoctorDocs

Hypothesis Generation

Jason Wright, PhD

Hypothesis Generation is a learning activity in which participants produce multiple competing hypotheses about a question, situation, or observation; examine each hypothesis against the available evidence; and refine, rank, or revise the hypotheses based on what the evidence supports. The deliberate production of multiple hypotheses is the method — a conversation that generates only one hypothesis and then tests it produces confirmation bias; a conversation that forces several hypotheses into the open before any is evaluated produces genuine inquiry. The activity's value depends on the discipline of holding off on commitment long enough for real alternatives to surface.

The approach draws on the scientific-inquiry tradition going back to Francis Bacon's Novum Organum (1620), which articulated the method of advancing understanding through the systematic testing of competing explanations. In education, Joseph Schwab's The Teaching of Science as Enquiry (1962) framed inquiry as the pedagogical counterpart to scientific hypothesis testing, and Jerome Bruner's 1961 article "The Act of Discovery" in Harvard Educational Review documented the learning benefits of participants generating their own hypotheses rather than receiving pre-formed answers. In medical education, Howard Barrows and Paul Feltovich's 1987 article on "The clinical reasoning process" in Medical Education showed how expert diagnosticians produce and test multiple hypotheses in parallel — the differential diagnosis — and how novices who commit to a single hypothesis too early reach wrong conclusions that better-calibrated reasoning would have caught.

Hypothesis Generation sits in the problem-and-inquiry family as the diagnostic-reasoning counterpart to Concept Mapping. Where Concept Mapping surfaces the structural relationships in a domain, Hypothesis Generation surfaces the possible explanations for something observed inside that domain. Both activities train a specific habit of mind: Concept Mapping trains the habit of articulating relationships; Hypothesis Generation trains the habit of producing alternatives before acting on any single one.

What It Is

Four structural components define Hypothesis Generation.

  • A question, situation, or observation that admits multiple explanations. The target is something where more than one hypothesis could plausibly account for the evidence — a cohort participant whose engagement is dropping, a workshop whose completion rate is lower than expected, a client whose progress has stalled. Questions with obvious answers produce performances of hypothesis generation, not real inquiry.
  • Explicit generation of multiple hypotheses. Participants are required to produce three to five hypotheses before any evaluation begins. The count is structural — single-hypothesis thinking is the diagnostic error the activity exists to counter, and asking for "some hypotheses" without a floor produces one or two and then rushes to testing. The generation discipline is the activity's load-bearing element.
  • Evidence examination against each hypothesis. Each hypothesis is tested against the available evidence. For each, two questions: what evidence would support this hypothesis? and what evidence would contradict it? The comparison across hypotheses is what produces the diagnostic calibration — hypotheses that nothing supports drop out, hypotheses that contradict clear evidence drop out, hypotheses that multiple evidence points support move up.
  • Refinement, ranking, or revision — not premature convergence. The close is not a vote on the right hypothesis. It is a refined view of the evidence-supported alternatives, often a ranked list with a leading hypothesis and one or two backups. Hypotheses that survive the evidence examination stay live. Converging on a single hypothesis before the evidence actually forces that convergence undoes the activity's mechanism.

When to Use It

Hypothesis Generation fits workshops where the target skill is diagnostic reasoning, alternative-thinking, or the discipline of holding off on commitment until the evidence actually supports it. It is especially strong when the cohort's work involves recurring diagnostic situations — coaching engagements where the real cause of a client's stuck-ness matters, program design where the real cause of low engagement matters, client work where the real cause of a slipping relationship matters.

Good candidates:

  • Diagnostic situations with multiple plausible explanations — engagement drops, performance plateaus, workshop outcomes that underperform expectations, client behavior changes.
  • Target skills that involve surfacing alternatives before acting — coaching, consulting, diagnostic work, any domain where premature commitment to a single explanation is a recurring failure pattern.
  • Cohorts of six to sixteen — enough to produce genuinely different hypotheses, small enough that every hypothesis gets examined in the time available.
  • Sessions of 60 to 90 minutes with time for individual generation, full-group share, evidence examination, and refinement.

Less suitable:

  • Questions with one obvious right answer — the activity produces a rehearsed performance of inquiry rather than real diagnostic work.
  • Emergencies or time-compressed decisions where acting on the first plausible hypothesis is actually correct — Hypothesis Generation is for situations where premature commitment is the failure mode, not situations where deliberation is.
  • Topics too abstract to admit concrete hypotheses — questions like "what makes a workshop effective?" produce hypotheses so general they cannot be tested against evidence.
  • Cohorts without enough domain foundation to generate hypotheses that differ meaningfully — all the hypotheses collapse into variants of the same generic explanation.

In-Person and Virtual Delivery

Hypothesis Generation runs cleanly in both in-person and virtual delivery. The activity is primarily verbal and analytical, and its core mechanism — forcing multiple hypotheses into the open before any is tested — works the same regardless of modality.

In person. Participants generate hypotheses on paper or sticky notes, which then get posted on a shared wall or board so every hypothesis is visible during the evidence-examination stage. The physicality of the posted hypotheses is an asset: moving hypotheses around as evidence changes their standing (keeping, refining, dropping, ranking) is tactile and clear. In-person delivery has a small advantage in group dynamics — the leader can read the room for participants who have converged prematurely and redirect them back to the alternatives.

Virtual. Participants generate hypotheses in a shared document or canvas, with each hypothesis visible as a labeled text block or card. The shared canvas persists through the session, and evidence examination happens by annotating each hypothesis with supporting and contradicting evidence. Virtual delivery's advantage is the written durability of the artifact — the hypothesis set and evidence annotations travel out of the session cleanly, which supports the post-session application where participants apply multi-hypothesis thinking to their own diagnostic work.

Choosing the modality. Neither modality is clearly stronger. In-person suits cohorts who benefit from tactile manipulation of hypotheses; virtual suits cohorts working in digital-first tools or multi-session programs where the hypothesis set needs to persist. For mixed-modality cohorts, run virtually with a shared canvas — the written hypotheses travel across the cohort regardless of where participants are physically.

How to Run It

Before the Session

  1. Select the target situation (leader design work). Choose a situation that admits multiple plausible explanations and has enough evidence available to actually test hypotheses against. Abstract situations produce abstract hypotheses; concrete situations with specific evidence produce usable diagnostic work.
  2. Prepare the evidence (leader design work). Assemble the evidence that will be examined — what participants know about the situation. For a cohort-engagement case, this might include attendance patterns, engagement signals, communication samples, pre-work completion data, peer observations. The evidence must be available before the hypothesis-examination stage, not invented during it.
  3. Optional — distribute background materials. For situations where context matters, send a short case brief a day or two before the session so participants arrive with the relevant context already loaded.

During the Session

  1. Present the situation and available evidence (8–10 minutes). Walk the cohort through the situation clearly. Share the evidence that's available. Flag any information gaps explicitly — what is not known, what would need to be discovered to fully test a hypothesis. Ensure every participant has the same evidence base before generation begins.
  2. Individual hypothesis generation (8–10 minutes). Each participant writes three to five distinct hypotheses that could account for the situation. The count is the discipline — require it explicitly and do not accept "I have two" as sufficient. Individual work before group discussion prevents early convergence on the first hypothesis voiced.
  3. Share hypotheses across the group (10–15 minutes). Every hypothesis goes up on the board or shared canvas, grouped loosely by similarity. Duplicates merge; distinct hypotheses stay separate. The group's full hypothesis set is visible before any evaluation begins.
  4. Examine each hypothesis against the evidence (15–25 minutes). For each hypothesis, two questions: what evidence would support this? and what evidence would contradict this? The group works through hypotheses in the order they appeared or in an order the leader chooses. Evidence-supported hypotheses get a mark; contradicted hypotheses get crossed out or revised; hypotheses with no evidence either way get flagged for what would need to be investigated.
  5. Rank or refine (10–15 minutes). The group produces a ranked list of the evidence-supported hypotheses — typically a leading hypothesis with one or two live backups. Participants who entered with one conviction often leave with two or three, which is the diagnostic calibration the activity is designed to build.

After the Session

  1. Apply multi-hypothesis thinking to participants' own diagnostic work (within one to two weeks). Each participant takes one situation from their own practice — a client they are diagnosing, a participant they are reading, an outcome they are trying to explain — and applies the same discipline: three to five hypotheses before evaluating any, evidence examination for each, ranked list at the close. Without this application, the activity builds the habit only inside the session.
  2. Bring the applied hypothesis work to the next session (in multi-session cohorts). Participants share one situation from their own practice, the hypotheses they generated, the evidence they examined, and the ranked view they landed on. The cohort's accumulated applications show how the discipline survives contact with real work.

An Example in Practice

A 75-minute workshop for a cohort of ten coaches on diagnosing cohort members who are falling behind. The workshop leader presents a specific case: a six-month cohort program, week eight of twelve, a specific participant named Sam. Sam attended all four of the first four weeks. Missed week five entirely. Present but quiet weeks six and seven. Has not submitted the last two pre-works. The coach has noticed and is trying to decide what to do.

The evidence, distributed before the session: Sam's initial application answers (highly motivated, specific outcome goals, clear willingness to commit the time); attendance and engagement pattern week by week; Sam's email reply latency (72 hours average, up from 18 hours in the first month); one peer observation from week seven ("Sam seemed distracted"); and one text Sam sent the coach on the morning of week five: "things are a lot right now." Participants arrive having read the case.

Individual hypothesis generation runs ten minutes. Every coach produces five hypotheses. The full-group share-out surfaces the set: time and workload overload, an emotional or life circumstance, a foundational skill gap Sam did not previously show, content-context mismatch (program shifted into material less relevant to Sam's situation), disengagement from the program's overall purpose, a relational rupture with another cohort member, accountability-structure mismatch (Sam works better with external deadlines than intrinsic ones).

Evidence examination runs twenty minutes. The "things are a lot right now" text directly supports the emotional-circumstance hypothesis and indirectly supports the time-workload hypothesis — both are live. Sam's initial application answers contradict the foundational-skill-gap hypothesis and the disengagement hypothesis; Sam entered motivated and skilled. The week-five inflection point (from strong engagement to absence) supports the emotional-circumstance hypothesis specifically — something happened in or around week five. Reply latency increase tracks the same timeline. The content-context mismatch and accountability-structure hypotheses have no specific evidence either way; flagged as requiring further inquiry.

Ranking runs twelve minutes. The cohort lands on a ranked list: emotional or life circumstance (leading hypothesis, multiple supporting evidence points); time and workload overload (secondary, partially supported); content-context mismatch and accountability-structure mismatch (live but unsupported, would require a direct conversation with Sam to test). The coaches notice something important: the intervention that follows depends heavily on which hypothesis is correct. A coach acting on the time-workload hypothesis would tighten the structure and add accountability; a coach acting on the emotional-circumstance hypothesis would reach out with care and make space. Acting on the wrong hypothesis would likely accelerate Sam's departure from the program rather than preventing it.

Variations

  • Silent Hypothesis Generation. All generation and initial evidence examination happens in writing on a shared document or board before any spoken discussion. Used when the cohort has dominant-voice dynamics that would compress the hypothesis set, or when the workshop leader wants to ensure every participant contributes hypotheses the group will actually see.
  • Iterative Hypothesis Generation. Evidence is released in waves — the initial situation, then more evidence at fixed points — with hypotheses refined or revised as each wave arrives. Forces participants to hold multiple live hypotheses longer and to notice which hypotheses survive new information. Used when the learning goal is specifically about updating on evidence rather than only about initial generation.
  • Structured-Differential Hypothesis Generation. Hypotheses must cover a named set of categories before the activity moves to evaluation. Borrowed from medical differential diagnosis, where hypotheses must span categories (infection, trauma, structural, systemic, psychological, etc.) rather than all coming from one category. Used when the cohort has a known bias toward one type of hypothesis and the workshop leader wants to force category coverage.
  • Expert-Panel Hypothesis Generation. Participants are assigned perspectives — generating hypotheses as specific named expert types (a coach, a therapist, a program designer, a data analyst). Each perspective tends to produce different hypotheses, and the comparison across perspectives surfaces hypotheses the cohort would not reach on its own. Used when the cohort's professional background is narrow enough that hypothesis generation would otherwise be homogeneous.

Which Method It Serves

Hypothesis Generation is primarily an Inquiry-Based Learning activity. The activity is inquiry in concentrated form — multiple hypotheses generated, tested against evidence, refined into a ranked view. The method activates every element of inquiry-based methodology: an open question at the center, hypothesis formation through participants' own generation, evidence examination through the testing stage, and synthesis through the refined ranking.

Hypothesis Generation also functions inside other methods. In Problem-Based Learning, hypothesis generation is often the opening move participants use when encountering an ill-structured problem — what are the possible explanations, and which ones does the available evidence support? In Case-Based Learning, hypothesis generation can precede the analytical framework application — generate hypotheses first, then work through them systematically. In Action Learning, hypothesis generation is a standard set-adviser move when a participant presents a stuck situation — the set helps the participant produce alternatives they had not considered before committing to a course of action.

Design Considerations

  • The discipline of generating multiple hypotheses before evaluating any determines the value of the activity. Hypothesis Generation is defined by the generation stage — producing three to five real hypotheses before any evidence examination begins. A cohort that generates two hypotheses and moves to testing has collapsed the activity into confirmation bias with a second opinion. A cohort that generates five hypotheses before anyone picks a favorite has actually done the inquiry work. Before running Hypothesis Generation, pressure-test the structure: is the minimum hypothesis count named explicitly (three to five), and is the generation stage protected long enough that participants actually reach that count? If either answer is soft, the activity's signature move — which is what separates inquiry from confirmation-driven analysis — will get lost.
  • Require individual generation before group discussion. The first hypothesis voiced in a group tends to anchor the group's subsequent generation. Individual work before share-out is what produces genuinely different hypotheses across participants rather than variations on the first one spoken.
  • Evidence must be available for real testing. A hypothesis that cannot be tested against any evidence is a guess, not a hypothesis. The pre-session evidence preparation is structural, not optional — if the activity is run on a situation without available evidence, the examination stage collapses into opinion exchange.
  • Hold off on ranking until the evidence has actually been examined. The ranking stage comes last for a reason. Cohorts that rank before testing produce rankings based on plausibility or familiarity, not evidence. The sequence — generate, test, then rank — is what produces evidence-calibrated rankings rather than intuition-calibrated ones.
  • Match the situation's complexity to the session's time. A simple situation with limited evidence can run in 60 minutes. A complex situation with substantial evidence needs 90 or longer — the examination stage alone can take twenty-five minutes if the evidence is rich. Under-scoping the time produces rushed testing and a ranking that misses calibration.

Where this goes next: Hypothesis Generation builds the discipline of producing multiple explanations before committing to any. The final activity in this family moves from diagnostic hypothesis work to structured design work — participants are given a real design constraint and generate, evaluate, and refine design solutions against it. See Design Challenge.

Read this next:

  • Design Challenge — the final activity in the problem-and-inquiry family; moves from diagnostic hypothesis work to structured generation and evaluation of design solutions.

Foundations of this activity:

  • Inquiry-Based Learning — the primary method Hypothesis Generation serves; the activity is inquiry in concentrated form.
  • Concept Mapping — the structural-visualization counterpart in this family; compare Concept Mapping's structural understanding with Hypothesis Generation's explanatory alternatives.
  • Case Analysis — the retrospective-analysis activity that Hypothesis Generation often precedes when a situation is being diagnosed before being analyzed.
  • The Guiding Principles of Active Learning — the principles Hypothesis Generation activates, especially honest reflection and status awareness.

Where this leads:

  • Which Active Learning Method Fits Your Workshop? — Hypothesis Generation is the strongest sub-activity when diagnostic reasoning or alternative-thinking is a target outcome.
  • Action Learning — the method where Hypothesis Generation most often appears as the set's opening move when a participant presents a stuck situation.

References

Barrows, H. S., & Feltovich, P. J. (1987). The clinical reasoning process. Medical Education, 21(2), 86–91.

Bruner, J. S. (1961). The act of discovery. Harvard Educational Review, 31, 21–32.

Schwab, J. J. (1962). The teaching of science as enquiry. Harvard University Press.

Frequently Asked Questions

What is Hypothesis Generation?

Hypothesis Generation is a learning activity in which participants produce multiple competing hypotheses about a question, situation, or observation; examine each hypothesis against the available evidence; and refine, rank, or revise the hypotheses based on what the evidence supports. The deliberate production of multiple hypotheses is the method — a conversation that generates only one hypothesis and then tests it produces confirmation bias; a conversation that forces several hypotheses into the open before any is evaluated produces genuine inquiry. The approach draws on the scientific-inquiry tradition and was specifically codified for medical education in Howard Barrows and Paul Feltovich's 1987 work on differential diagnosis.

When should I use Hypothesis Generation?

Hypothesis Generation fits workshops where the outcome is diagnostic reasoning or investigative judgment — the ability to consider alternatives before committing to one. Especially strong for coaches, consultants, and creators who need to diagnose complex situations in their own work (why is this client stalling, what's actually causing this offer to not convert, why is this cohort disengaging). Less suitable for content with clear right answers, for outcomes that don't require holding alternatives in mind, or for sessions too short to surface multiple genuine hypotheses before evaluation begins. The discipline of holding off on commitment long enough for real alternatives to surface is what the activity exists to build.

Why generate multiple hypotheses if only one is going to be right?

Because the cognitive lift comes from holding alternatives in mind before committing, not from arriving at the right answer fast. Howard Barrows and Paul Feltovich's 1987 research on clinical reasoning showed that expert diagnosticians produce multiple hypotheses in parallel — the *differential diagnosis* — while novices commit to a single hypothesis too early and reach wrong conclusions that better-calibrated reasoning would have caught. Forcing several hypotheses into the open before any is evaluated is what produces genuine inquiry; jumping to one and testing it produces confirmation bias. The discipline transfers — a participant who has practiced producing multiple hypotheses in workshop scenarios is more likely to produce them in their own work afterward.

Tagged

Free resource

See where your workshops stand

The Workshop Health Check pinpoints exactly where your workshops are losing impact — and gives you a clear path forward.

Take the Health Check →