DesktopConsultancy

Same Persona, Seven Mirrors: Persona Validation

17 December 2025

We've conducted a comprehensive validation study of our AI personas to ensure they are realistic, coherent, and useful for market research. This study uses the latest LLMmodels to assess the internal consistency and realism of our personas, ensuring they represent believable, coherent human beings.

Persona Validation

Personas are not clipart

Most “AI personas” on the internet are basically clipart. A demographic label, a job title, a couple of quirks, then the model does the rest. Sometimes it works. Often it doesn’t. It drifts, it stereotypes, it fills gaps with whatever the training data thinks is most likely.

At Mass, we treat personas differently. A persona is a research instrument. If it cannot stay itself across repeated questioning, it is not a respondent, it’s a text generator wearing a name badge.

That is why we validate. Not because we love process, and we do, but because “generic” personas fail in predictable ways: they collapse into archetypes, they over-index on demographic markers, and they look coherent until you put pressure on them.

This is about what we learned from putting our personas under that pressure. Persona design, prompt constraints, temperature, context depth, and the specific model you pick all matter. Models are not interchangeable. Pretend they are, and your panel will look fine right up until you ask anything subjective, identity-adjacent, or culturally loaded, then it starts free-associating its way into stereotypes.

We tested a range of leading models across providers under the same conditions. Same personas, same questions, same set-up. What changed was the failure mode. The differences weren’t subtle. They were philosophical.

Same people, same questions, lots of chances to drift

We ran repeated generations across a grid of conditions designed to stress the system in the places it usually breaks:

  • Two deliberately asymmetric personas: one profession-anchored (stable craft identity), one identity-anchored (dual heritage, faith, lived experience, and social context)
  • A mix of simple preference questions and applied research questions (product choice, sustainability, decision trade-offs)
  • Multiple temperature settings (to test bounded variation vs drift)
  • Multiple levels of persona context depth (to test how much scaffolding is required)

We then audited outputs for consistency, cultural coherence, stereotyping, contradictions, leakage between personas, and dialect caricature. Quantitative scoring is useful. Qualitative failure modes are the difference between “high consistency” and “high consistency in the wrong direction”.

On the quantitative side, we deliberately avoided “did it say the exact same sentence” as a metric, because that’s how you end up rewarding robotic repetition. Instead, we scored whether responses stayed semantically aligned across repeat runs:

  • Semantic consistency: embed each response into a vector representation of meaning, then compare runs using cosine similarity (same idea, different words should still count as consistent)
  • Response length stability: does the persona keep a stable level of verbosity, or swing wildly depending on mood and sampling
  • Lexical drift: does the persona keep a coherent voice, or does it swap registers and vocabulary as if it’s a different person each time

What we found: the patterns

1) The Identity Complexity Gap, or: jobs are easy, people are hard

Across models, the profession-anchored persona was easier to keep stable than the identity-anchored persona. Not because identity is “too complex” in a hand-wavy way, but because the models have strong priors for professional archetypes. A job title has countless training-data grooves: job descriptions, portfolios, forum posts, tool arguments. A culturally situated person has fewer grooves, and the grooves that exist are more likely to be cultural shorthand.

In practice, that gap shows up as one persona becoming stable to the point of overfitting, while the other becomes unstable, stereotyped, or both.

If you want to do identity-led research, model choice is not a technical detail. It is the research design.

2) Demographic anchoring, or: the postcard problem

When persona context doesn’t explicitly answer a subjective question, models fill the gap. They do not fill it neutrally. They fill it with high-probability associations from training data.

That’s demographic anchoring: a demographic marker becomes an autocomplete trigger.

It shows up as “helpful” cultural objects getting stapled to the person:

  • A country becomes an animal.
  • A place becomes a landscape.
  • A faith becomes a vibe, or worse, gets ignored when it conflicts with a cosy Western trope.

Technically, this is just probability. When context is sparse, the model samples from the highest-likelihood continuations. Cultural stereotypes are high-likelihood continuations because training data is full of them. The implication is uncomfortable but practical: thin personas are not just lower quality. They are higher bias risk.

3) Consistency vs creativity, or: temperature is a character stability dial

Lower temperature tightens sampling. You get repeatability, and sometimes hyper-consistency: the persona becomes a loop.

Higher temperature loosens sampling. You get variety, and sometimes drift: the persona becomes a series of plausible strangers.

Different models break differently. Some become brittle and repetitive at lower temperature. Some stay improvisational even when you try to clamp them down. Some achieve what you actually want: bounded stochasticity, where the person varies naturally but stays within a coherent cluster of truths.

The practical point is not “always use a low temperature”. The point is that temperature interacts with model tuning. You have to test, not assume.

4) Context depth is not fluff, it’s scaffolding

More persona context generally improved results. Not equally for every model, and not in exactly the same way, but the direction held.

More scaffolding gives the model more anchor points. It reduces the chance that it will reach for demographic shortcuts because there are other salient details to pull from: routines, money habits, family structure, constraints, values that actually bite.

If you are building identity-anchored personas and you are not paying for context depth, you are essentially asking the model to invent a person from a handful of demographic tags. It will. You just won’t like what it invents.

5) Failure modes that matter in real research

The failures aren’t abstract. They map directly onto ways panels can mislead you:

  • Hyper-consistency: semantically similar answers that are basically the same response key every time. Looks stable. Feels fake.
  • Lack of state: basic facts and preferences mutate across repeated asks. Looks lively. Breaks longitudinal use.
  • Response leakage: one persona borrowing another persona’s tells. Disqualifying for panels unless you can guarantee isolation.
  • Religious and cultural contradictions: the model resolves messy inputs by walking straight into a coherence or respect failure.
  • Dialect caricature: a “soft lilt” becomes phonetic theatre. This is a brand risk, not a stylistic quirk.

What this means

If you want synthetic respondents you can trust, you have to stop treating personas as content and start treating them as research instruments.

That means:

  • You validate against the model you will actually run. “We can swap providers later” is not a plan. Model choice changes the priors. The priors change the people.
  • You scope questions to what the persona can credibly answer. Values and trade-offs tend to be more stable than “favourite animal” style questions, which expose priors fast.
  • You budget for context depth where identity is doing work. Thin identity personas are stereotype volatility with a name tag.
  • You actively test for contamination and drift. If the same person can’t answer the same question in the same semantic neighbourhood, you do not have a respondent.

What Mass does differently

Synthetic panels can work. But if you don’t validate them, you are not doing research. You are doing improv with a spreadsheet.

Mass exists because “make a persona” is not the hard part. Making a persona you can trust in a research context is.

We build personas as research tools: grounded, constraint-aware, and designed to behave consistently under repeated questioning. Then we validate them against real model behaviour, not wishful thinking, because the only honest question is: does this persona stay itself when you ask again, and does it stay culturally coherent when the prompt leaves gaps?

If you want the full technical report, with per-model variance, failure-mode concentration, and the methodology in detail, enter your details below and we'll reach out for a demo.

Ready to work with validated, consistent personas?

Experience our validated persona system designed for research-grade consistency. Our personas maintain semantic coherence, avoid stereotyping, and stay true to their identity across repeated questioning, ensuring reliable, trustworthy insights.

  • Consultancy (SRaaS): research design, audience and scenario set-up, running the simulation, and interpreting the results with you
  • Custom deployments and dedicated environments for high-volume or privacy-sensitive data
  • Custom audiences, grounding in your own data, and CRM integration
  • Comprehensive API access to integrate Mass into your own workflows
  • Reduced pricing for non-profit organisations, academic institutions, and public-interest work

By clicking the "Submit" button above, you acknowledge that Mass may use the information you provide to contact you about Mass's products and services and you agree to be contacted by Mass in accordance with Mass's Privacy Policy.