Methodology
Every published experiment on this site follows the same method. This page explains it once, so each experiment page can stay focused on its own result.
What an experiment is
An experiment is a prompt — or a small set of prompts — run against one or more AI models, with the variables that might change the answer held explicit: which model, which provider, whether web search (grounding) is on, and which memory profile (if any) conditions the assistant beforehand. Each of those combinations is a variant. An experiment compares variants against each other, not a single run against nothing.
Why repeat-sampling
A single call to an AI model is one draw from a distribution, not a fact. Ask the same question of the same model twice and you can get two different citations, or a brand mentioned in one answer and not the other. A finding built on one sample is an anecdote.
So every variant in an experiment is sampled multiple times — the "repeats" you'll see in each experiment's method table. That turns "the model said X" into "the model said X in 4 of 5 samples," which is a claim you can actually trust, and a number you can quote. The consistency headline and the consensus-cited-sources counts on every experiment page come directly from this — they're never a read of one lucky (or unlucky) response.
Realistic memory, not persona-stuffed prompts
Real people do not type persona descriptions into ChatGPT. They type "best email signature software" — a few words, the way anyone actually asks — and the assistant already has some notion of who's asking from memory and prior conversation, the same way it would for a real returning user.
Some AEO tools simulate a persona a different way: they concatenate demographic and firmographic attributes into first-person prose and prepend the whole block to the query itself. A verbatim example we've seen from a market-leading tool (July 2026):
"i am a 25-34, 35-44, or 45-54 year old marketing manager in the finance, healthcare, manufacturing, or pharma industry. my job seniority is at the director, manager, senior, or vp level… email signature management software."
We want to frame this as a method difference, not a takedown — how you simulate a persona changes your results, and the data speaks for itself once you look at it side by side. Two things follow from prompt-stuffing that are worth being clear-eyed about:
- No human types this. The results measure how a model responds to a query that doesn't exist in the real world. The persona contaminates the query instead of conditioning the context around it.
- The prompts leak. Grounded models turn a query like the one above into a real web search — and that search can surface in the target site's own Google Search Console, as a literal query string nobody typed. A tool built to measure AI visibility can end up polluting the search data of the very sites it's measuring for.
AEO Mission Control conditions the assistant's memory instead — a persona profile the model is given as context — and keeps the query itself short and natural, the way someone would actually type it. Every published experiment shows exactly which memory profile (if any) was used, in full, under "Memory profile" — nothing about how the persona was built is hidden behind the result.
Reading a result
Each experiment page shows, per variant: a consistency headline over every sample, the sources the model actually cited (with how many of the samples cited them), and the full text of the most representative sample — the "medoid," the one response whose content is most typical of the whole batch, not cherry-picked and not the first one that happened to run.
Every experiment is published in full — the prompts, the memory profile, and every variant's method and results. Nothing here is trimmed for a good headline.