DesktopConsultancy

Dead Eyes, Melted Tools, and the One Thing You Cannot Fake

21 December 2025

We stress-tested 14 image models for persona profile photos. The results were not "roughly the same". They were different species.

Persona Image Generation Testing

Who cares what a face looks like

Brands come to us to test ideas with hundreds of diverse AI personas instantly. Those personas are synthetic research participants, but the work only lands if they feel like participants. The profile image is one of those credibility checks. If the face reads as “AI”, the feedback reads as “toy”. If it reads as a real person, the platform reads as research.

One model gives you a believable 58-year-old craftsperson in a Scottish workshop, weathered skin, proper mood, a glint in the eyes. Another gives you the same man with a sweater zipper that turns into “a distorted, melted blob of metal” and chisels that look like they were generated by a haunted label maker. That’s not a minor variance. That’s the difference between “I trust this research” and “this is obviously fake”.

We tested 14 image models across 4 personas (56 images total), rated each image across realism, style, authenticity, technical quality, persona match, and overall score, similar to how we test and score all our personas.

The short version of what we learned

The good news is that several models already produce images we would happily put in inside Mass. The more honest news is that most of the failures live in the same places, and they are the places humans notice first: eyes, text, and anything that requires basic geometry.

If you only take one thing from this test, take this: image generation is not “pick a model and forget about it”. It’s a product surface. We have to treat it like one.

1) Realism is a chain, not a score

Most models can ship a convincing thumbnail. Realism breaks when you zoom, and the breakpoints are surprisingly consistent: skin microtexture, hairline transitions, object geometry, and lighting that actually obeys physics.

The strongest cluster across our set was steady: flux-2-max, flux-2/lore-realism, gpt-image-1.5, seedream/v4.5, and often nano-banana-pro. On multiple personas, they landed in the “looks real at a glance” bracket.

Stellan, the Scottish workshop craftsperson, is a good example of why this is a chain. Several models got the face right, then quietly melted the tool rack into modern sculpture. By contrast, omnigen-v2 was consistently less convincing across personas, with the classic smooth skin and flat light that reads synthetic immediately.

2) Eyes decide whether the persona feels real

Eyes are the trust surface. You can forgive a slightly mushy background. You cannot forgive catchlights that do not match the scene, pupils that are not quite circular, or that oddly vacant “painted” iris texture.

Even strong images showed eye tells, though very hard to notice in the context of a persona. gpt-image-1.5 sometimes fell into the physically impossible reflection problem, one window in one eye, a different universe in the other. Seedream could deliver excellent mood and skin, then give you a pupil edge that looks slightly serrated. nano-banana-pro occasionally “solves” the problem socially by squinting, which works, but it is also a tell.

3) Background text is still radioactive

Persona images live in UI, and UI has text. The problem is that background text is still one of the loudest “this is AI” alarms.

Zola’s bookshelf scene made this obvious. Close enough to be funny, wrong enough to undermine trust. This is not a reason to avoid images, it’s a reason to treat text as a controlled ingredient. We will either avoid prompts that invite readable text, or deliberately post-process it (inpainting, compositing, or clean overlays).

4) "Candid smartphone" is a genre, and models default to cinema

A lot of models want to give you a polished portrait, because that's what the internet has rewarded for years. We asked for candid smartphone photos, unposed, slightly imperfect, a person who feels like they have a life outside the frame.

Some models can do the "phone portrait mode" vibe, others drift into cinematic headshot territory. Seedream in particular produces oddly beautiful people, so beautiful we had to add imperfections to the prompt to make them feel more real. The fix here is mostly product discipline: prompt constraints that punish studio lighting and reward "real phone photo" cues, plus model selection that matches the aesthetic we actually want.

5) Persona match is strong, detail match is fragile

Overall, persona match was the highest-scoring dimension in our dataset. Models can generally read a description and cast a plausible person. The fragility shows up in the specific details that make a persona feel lived-in rather than generic.

Multiple models correctly included “jar on shelf”, then invented the wrong contents. Joris was meant to be 32, but some models nudged him older. These are small drifts, but in a research product they create cognitive dissonance. The fix is twofold: tighter prompt constraints on age and physical details, and routing to models that behave better on fidelity.

What this means for Mass (and what we’ll do next)

This experiment wasn’t about dunking on models. It was about choosing what belongs in production, and building guardrails around what still breaks.

Here’s the forward plan:

  • Shortlist for production: based on this round, the current front-runners for Mass persona images are flux-2-max, flux-2/lore-realism, seedream/v4.5, seedream/v4, gpt-image-1.5, and nano-banana-pro. They are not perfect, but they feel real in different ways.
  • Model routing rather than one “winner”: we will route models based on persona, scene complexity, and risk. A workshop with tools is a different job to a bright outdoor portrait.
  • Automated quality gates: eye checks, background text checks, and a small set of “do not ship” heuristics. If an image fails, we re-roll or switch models. This is much cheaper than rebuilding trust after a brand has already seen a cursed bookshelf.
  • Prompt constraints that match the product: “candid smartphone” needs to be enforced, not politely requested.
  • Continuous evaluation: the judge loop stays. Every new model change gets measured against a fixed set of personas, so we do not quietly regress.

The upbeat part is simple: we now have evidence, not vibes. The path to better persona imagery is not a single breakthrough. It’s selecting reliable models, then systematically removing the failure modes that users actually notice. This might seem overkill for a platform like Mass, but for us, the details matter.

Check out the resulting images below.

Model
Stellan Novak
Zola Pereira
Joris Alston
Daria Eldoria
flux-2-max
flux-2-max
flux-2-max
flux-2-max
flux-2-max
flux-2/lore-realism
flux-2/lore-realism
flux-2/lore-realism
flux-2/lore-realism
flux-2/lore-realism
gpt-image-1.5
gpt-image-1.5
gpt-image-1.5
gpt-image-1.5
gpt-image-1.5
seedream/v4.5
seedream/v4.5
seedream/v4.5
seedream/v4.5
seedream/v4.5
seedream/v4
seedream/v4
seedream/v4
seedream/v4
seedream/v4
nano-banana
nano-banana
nano-banana
nano-banana
nano-banana
nano-banana-pro
nano-banana-pro
nano-banana-pro
nano-banana-pro
nano-banana-pro
z-image/turbo
z-image/turbo
z-image/turbo
z-image/turbo
z-image/turbo
imagen4
imagen4
imagen4
imagen4
imagen4
hunyuan-image/v3
hunyuan-image/v3
hunyuan-image/v3
hunyuan-image/v3
hunyuan-image/v3
reve
reve
reve
reve
reve
longcat-image
longcat-image
longcat-image
longcat-image
longcat-image
sana/v1.5
sana/v1.5
sana/v1.5
sana/v1.5
sana/v1.5
omnigen-v2
omnigen-v2
omnigen-v2
omnigen-v2
omnigen-v2

Ready to work with validated, consistent personas?

Experience our validated persona system designed for research-grade consistency. Our personas maintain semantic coherence, avoid stereotyping, and stay true to their identity across repeated questioning, ensuring reliable, trustworthy insights.

  • Consultancy (SRaaS): research design, audience and scenario set-up, running the simulation, and interpreting the results with you
  • Custom deployments and dedicated environments for high-volume or privacy-sensitive data
  • Custom audiences, grounding in your own data, and CRM integration
  • Comprehensive API access to integrate Mass into your own workflows
  • Reduced pricing for non-profit organisations, academic institutions, and public-interest work

By clicking the "Submit" button above, you acknowledge that Mass may use the information you provide to contact you about Mass's products and services and you agree to be contacted by Mass in accordance with Mass's Privacy Policy.

Explore more insights

New Models, Same Hive

Aug 7, 2026

We have quietly re-housed the swarm. Mass now runs on the latest generation from Google, with a sharper Pro tier for the heavy lifting. For certain clients, we also route work through the newest Chinese models, because when you care about prediction, you go where the calibration is.

Features

Same Persona, Seven Mirrors: Persona Validation

Dec 17, 2025

We've conducted a comprehensive validation study of our AI personas to ensure they are realistic, coherent, and useful for market research. This study uses the latest LLMmodels to assess the internal consistency and realism of our personas, ensuring they represent believable, coherent human beings.

Research

The Reality Gap: Decoding the 2025 Christmas Ad Season

Dec 15, 2025

Analysing how specific slices of British society respond to Christmas adverts, using Mass Vision's persona-based sentiment analysis engine.

Research

The Science Behind Mass

Oct 24, 2025

The evolution of market research is witnessing a paradigm shift from traditional human-centered methodologies to AI-powered persona simulations. Here we examine the scientific foundations supporting the persona approach for Mass and demonstrates why AI personas offer advantages over conventional focus group methodologies for initial market validation and feedback generation.

Research

Meet Our Personas

Oct 20, 2025

Discover our diverse collection of AI personas representing real people from around the world. Each persona brings unique perspectives, backgrounds, and insights to help you understand your audience better.

Features