Find themes in survey free text without cherry-picking

For: Insights analyst or customer research lead with thousands of open-text answers

Pattern: TournamentNeeds scaleDesigned for 30 to 600 agents

The pain today

Open-text answers get a skim and a word cloud. Themes are chosen by whoever reads first, quotes are cherry-picked, and small themes that matter vanish.

The ask

I attached the open-text answers from our customer survey with respondent details removed. Propose candidate themes, test each against the answers, keep the ones that hold, and show me supporting and contradicting quotes for each, including small themes.

Plain words, as you would say it to a colleague. Edit it to fit your case before you send it.

What you attach or connect

  • Open-text answers with respondent details removed
  • Survey questions
  • Segment labels, if any, at group level only

The unit of work

One worker task per one batch of answers per candidate theme.

Why a swarm fits

Candidate themes are cheap to propose from small batches, and each can be tested against other batches independently. Rounds of independent judging remove duplicates and themes that only one batch supports.

Not for

Measuring how common a theme is with statistical confidence. That needs a coded sample and an analyst.

The decision tree

6 typed decisions, each with an action for every answer

At fixed moments in a run, the engine puts one narrow question to a decision model. The decision model never writes text: it answers yes or no with a probability, picks from listed options, or gives a score, about a small slice of the material. The engine then does exactly what this tree says, which is what makes the run auditable. The thresholds are the template's design values, not measured results.

  1. Planner, while planning

    Scope checkYes or no, with a probability

    Before work starts on a unit

    Does this open-text answer say something about the respondent's experience, rather than being blank, a placeholder or off topic?

    Sees only: One open-text answer and the survey question it replies to

    Why: Empty answers are dropped before any worker spends effort on them.

    • Yes: 0.60 or higherthenAccept
    • Unsure: 0.30 up to 0.60thenAccept
    • No: below 0.30thenSkip this unit
  2. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Does this verbatim answer express the candidate theme as it is worded, rather than a neighbouring complaint or a different product area?

    Sees only: One candidate theme's wording and one verbatim answer offered for it

    Why: Themes stand on their quotes, so a stretched quote is refused.

    • Yes: 0.85 or higherthenAccept
    • Unsure: 0.50 up to 0.85thenMark unresolved
    • No: below 0.50thenReject and retry
  3. Run control, between rounds

    Another round?A score

    Between rounds

    How well do the verified quotes from batches this theme was not proposed from support it?

    Sees only: One theme and its verified quotes from unseen batches only

    Why: A theme advances only when answers it has never seen also say it; cherry-picked themes stop here.

    • High: 0.70 or higherthenContinue
    • Middle: 0.40 up to 0.70thenContinue
    • Low: below 0.40thenStop
  4. Reconciler, while merging

    Conflict checkA choice among options

    While reconciling

    How do these two surviving themes relate to each other?

    Sees only: Two theme wordings with a few verified quotes each

    Why: Removes duplicates without erasing a small theme that only looks like a large one.

    • Same theme, mergethenAccept
    • Distinct themesthenContinue
    • One contains the otherthenEscalate to a strong model
    • They contradict each otherthenMark unresolved
  5. After workers, the judge checks

    Evidence checkYes or no, with a probability

    After a worker answers

    Is this quote free of any name, job title, location or event specific enough to identify the respondent or another person?

    Sees only: One verbatim quote chosen for the report

    Why: Identifying quotes are kept out of the report even when they are the most vivid.

    • Yes: 0.90 or higherthenAccept
    • Unsure: 0.60 up to 0.90thenAsk a person
    • No: below 0.60thenSkip this unit
  6. Accountable person, before anything is settled

    Person decidesYes or no, with a probability

    Before anything is reported as settled

    Does the surviving theme set include small themes, contradicting quotes or merged themes that an analyst should confirm before it is shared?

    Sees only: The surviving theme list with quote counts by round and the merge log

    Why: The insights lead approves what the organisation will treat as the voice of its customers.

    Accountable: The insights lead approves the theme set and checks that no quote identifies a respondent.

    • Yes: 0.30 or higherthenAsk a person
    • Unsure: 0.10 up to 0.30thenAsk a person
    • No: below 0.10thenAccept

The fleet: who does what

Model tiers by role, not brands: you choose the models. Strong reasoning models plan and reconcile, small fast models do the wide work, and the judge is a decision model from a different family, so it does not share the workers' blind spots.

  1. Planner

    A strong reasoning model batches the answers and sets the rules a theme must meet to advance.

    Decisions here:1. Scope check

  2. Workers

    Small fast workers from an open-weight family propose themes per batch, then test surviving themes on unseen batches.

    Designed for 30 to 600 agents, one worker task per one batch of answers per candidate theme. Each worker receives only its own unit.

  3. Judge, from a different model family

    Decision models from a different family score each theme per round: distinct, supported by quotes, not a duplicate.

    Decisions here:2. Evidence check5. Evidence check

  4. Reconciler

    A strong reasoning model merges near-duplicates and reports surviving themes with quotes for and against.

    Decisions here:3. Another round?4. Conflict check

  5. Accountable person

    The insights lead approves the theme set and checks that no quote identifies a respondent.

    Decisions here:6. Person decides

Checked before anything is accepted

  • A theme advances only if quotes from batches it was not proposed from support it
  • Every quote is verbatim and traceable to an answer id
  • Contradicting quotes are collected for each surviving theme
  • Themes dropped in each round are kept in a log with the reason

What comes back

  • Surviving themes with supporting and contradicting quotes
  • Small themes kept because the quotes are strong
  • Dropped and merged themes, with reasons
  • Answers that fit no theme

What to measure

  • Agreement between the theme set and an analyst-coded sample
  • Share of quotes the analyst judges correctly assigned
  • Analyst hours per survey wave
  • Cost per thousand answers read

Names of measures only. No result is claimed for this template.

Templates open in the workspace chat with the ask filled in. Nothing runs until you send it.

Get early accessSign in to use

Find conflicting definitions across the data dictionary

For: Data governance lead or analytics engineering manager

Active customer, net revenue and churn are each defined several ways across tables, models and dashboards.

Pattern: Map, verify, reduceNeeds scale6 decisionsDesigned for 40 to 1,500 agents

Cross-check a source-to-target mapping before migration

For: Data migration lead or solution architect moving a legacy system to a new platform

The mapping sheet has thousands of fields, filled in by different people from column names.

Pattern: Cross-examinationNeeds scale6 decisionsDesigned for 40 to 2,000 agents