New Models, Same Hive
We have quietly re-housed the swarm. Mass now runs on the latest generation from Google, with a sharper Pro tier for the heavy lifting. For certain clients, we also route work through the newest Chinese models, because when you care about prediction, you go where the calibration is.

Models age in dog years
In AI product work, six months is a career. A model that felt sharp in December becomes a redirect target by summer. Retirements land quietly. Defaults rot louder.
Mass runs personas, reports, and workflows on real model APIs, not a marketing slide that says “powered by AI”. When the underlying models move, we move with them. Leave the defaults on last season's preview IDs and you get quieter failures: more hedging, more tokens spent saying less, more drift when you ask the same person the same question twice.
So we upgraded the stack. Not because shiny version numbers look good in a changelog, but because research-grade synthetic respondents deserve a current brain.
What changed
The everyday workhorse got an upgrade
Persona chat, report responses, questionnaires, debates, feedback loops, and workflow steps now run on the latest stable Flash-class release from Google. Multimodal, thinking-capable, tuned for agent loops, and notably less verbose than the previous generation.
That matters more than it sounds. Personas should sound like people, not like a committee drafting a press release. The newer generation wastes fewer tokens on throat-clearing and hedging. It gets to the point faster, holds structured output more reliably, and keeps tool use tighter when a workflow has several steps chained together. Shorter, cleaner generations also mean lower cost per research turn, without pretending that cheaper always equals worse.
The heavy lifts moved to the new Pro tier
Unified reports, questionnaire synthesis, and idea-generation passes now sit on the current Pro line. Those jobs need sustained reasoning and long-document coherence more than raw speed. Flash handles conversation and volume. Pro handles the documents that have to hold together under scrutiny.
The technical difference is not just “bigger”. The newer Pro stack holds a consistent argument across thousands of tokens, contradicts itself less halfway through a report, and follows a schema instead of inventing its own structure when the prompt gets complicated. If you have ever read an AI-generated report that starts confident and ends in a different universe, that is the failure mode this tier exists to avoid.
A cleaner menu
The picker now surfaces the current generation. Older preview IDs remain in the system for compatibility, but they are no longer offered as first-class choices. Saved workflows keep whatever model they were built with until you change them. New work starts on current metal.
The Chinese models, and why prediction people care
For certain clients and certain workloads, we also route through the latest Chinese frontier models. The point is not the badge. The point is what they are good at.
The newest Chinese releases are unusually strong at prediction-shaped tasks. Part of that is training emphasis: a lot of the recent work coming out of that ecosystem is optimised for reasoning over numbers, trends, and structured signals rather than open-ended chat. Part of it is calibration. Ask for a forecast, a likelihood, or a ranked set of outcomes, and the better models in that family commit to a distribution instead of performing confidence and then waffling. They are also very good at pattern completion over dense, messy inputs, which is exactly what market signals look like before anyone has cleaned them up for a slide deck.
For persona work, that shows up as more grounded answers to questions like “what would this segment actually do”, not just “what would they say”. For report work, it shows up as forecasts that behave like forecasts. We use them where that behaviour is the point.
Why this is not just plumbing
Model choice is research design. We have written about that before. Different models fail differently: one overfits to job titles, another fills identity gaps with postcard stereotypes, another becomes a loop at low temperature.
The questions you ask a persona about price sensitivity, brand trust, or cultural fit inherit the priors of the model underneath. Run those questions on a retired preview because nobody rotated the constants, and the panel goes stale while the UI still looks polished.
The bees in the hero image are not subtle. A hive only works when the workers are current. Swap the wrong cohort in, and the honey tastes like last quarter's training cut-off.
What you should notice
If you use Mass regularly, expect:
- Tighter persona replies on the conversational paths. Less padding, same intent.
- Stronger long-form structure on the Pro-backed reports.
- A model list that matches what providers actually ship in 2026.
If nothing feels dramatically different, that is the point. The instrument should feel continuous. The internals should not.
The boring moral
Frontier models move. Product defaults should move with them, on purpose, with the research consequences named out loud.
We will keep rotating the hive as providers ship stable replacements. Not every week. Not for the thrill of the changelog. Often enough that when you ask a Mass persona a hard question, you are talking to a living model, not a redirect with a nostalgic name.
Want to run research on the latest models?
Mass keeps the model stack current so your personas do not answer like it is last quarter. Book a demo and hear what your audiences sound like on a fresh brain.
- Consultancy (SRaaS): research design, audience and scenario set-up, running the simulation, and interpreting the results with you
- Custom deployments and dedicated environments for high-volume or privacy-sensitive data
- Custom audiences, grounding in your own data, and CRM integration
- Comprehensive API access to integrate Mass into your own workflows
- Reduced pricing for non-profit organisations, academic institutions, and public-interest work
Explore more insights
Dead Eyes, Melted Tools, and the One Thing You Cannot Fake
Dec 21, 2025
We stress-tested 14 image models for persona profile photos. The results were not "roughly the same". They were different species.
Same Persona, Seven Mirrors: Persona Validation
Dec 17, 2025
We've conducted a comprehensive validation study of our AI personas to ensure they are realistic, coherent, and useful for market research. This study uses the latest LLMmodels to assess the internal consistency and realism of our personas, ensuring they represent believable, coherent human beings.
The Reality Gap: Decoding the 2025 Christmas Ad Season
Dec 15, 2025
Analysing how specific slices of British society respond to Christmas adverts, using Mass Vision's persona-based sentiment analysis engine.
The Science Behind Mass
Oct 24, 2025
The evolution of market research is witnessing a paradigm shift from traditional human-centered methodologies to AI-powered persona simulations. Here we examine the scientific foundations supporting the persona approach for Mass and demonstrates why AI personas offer advantages over conventional focus group methodologies for initial market validation and feedback generation.
Meet Our Personas
Oct 20, 2025
Discover our diverse collection of AI personas representing real people from around the world. Each persona brings unique perspectives, backgrounds, and insights to help you understand your audience better.