The questions
For every product, Aethos generates realistic buyer questions across seven commerce intents — discovery, comparison, budget, occasion, problem-solving, gifting and trust. They read like real ChatGPT asks (“Are expensive cashmere sweaters actually worth it, or are mid-range brands just as good?”), and they never contain your product name — a prompt that names your product would trick the engine into echoing it and inflate your score. Mentions have to be earned.
The sampling
Live scans ask an answer engine — with web search — a rotating window of those questions per product per scan, and the window moves each week so successive scans cover different questions. Each weekly scan samples 2 questions per product, each asked once, on every one of the 3 answer engines — so a weekly figure is a monitoring sample, not a verification. We never report that an answer changed on the strength of one sample. Every scan surface tells you the runs and questions behind the number it shows, measured from the answers themselves rather than from a setting. Products are taken stalest first, and if a catalog is larger than the week's allowance the scan tells you exactly how many products it covered and the rest lead the next weekly rotation. Manual re-scans ask each question once and are throttled to one per product every 10 minutes, so costs stay predictable.
How much of your catalog we refresh, and how often
A product costs the same every scan: 2 questions × 1 run × 3 engines, plus 1 call to classify the tone of what came back — 7 engine calls in total. Your plan buys a fixed call budget each week, so what it really buys is a number of products:
- Starter — 15 products refreshed every week (105 engine calls), out of up to 250 tracked, so your whole catalog is refreshed every 17 weeks.
- Growth — 40 products refreshed every week (280 engine calls), out of up to 800 tracked, so your whole catalog is refreshed every 20 weeks.
Products are always taken stalest first, so the oldest score is the next one bought. If your catalog is larger than the number your plan tracks, the scan still runs — it covers the stalest products up to that number and tells you how many were left out, rather than skipping your store.
Adding an answer engine raises the number of calls a product costs. That raises what your plan spends; it does not shrink what your plan promises. We would rather tell you the arithmetic than quietly shrink your coverage — and the page you are reading is generated from the same numbers the scanner spends, so it cannot say one thing while the scanner does another.
How several engines combine
Every engine your deployment has switched on answers the same questions, and each engine is weighted equally in your score and share of voice — never by how many answers it happened to produce. An engine that failed half its calls would otherwise quietly count for less than one that answered everything, which would make your headline number a fact about our infrastructure rather than about you. So each engine's own presence-rate is measured first, and the combined figure is the average of those.
And the combined figure never appears on its own. Wherever you see a blended number — your visibility score, a product's score, share of voice, the report you can forward — the per-engine parts sit beside it, because “55% overall” means something very different when it is ChatGPT 80% and Gemini 30% than when both are 55%. The same rule applies to sampling depth: run counts are counted per engine and are never added together, so a question put to 3 engines is 1 run on each of them, never one number 3 times larger.
One exception, and it exists to protect you from us: an engine that completed fewer than two calls in a scan is shown with its rate but left OUT of the combined number. Equal weighting without that floor would let a single lucky answer from a rate-limited engine carry the same weight as another engine's dozen and move your headline by thirty points. Under-sampled engines are labelled wherever they appear, never quietly dropped — and if no engine reaches the floor we pool every answer instead, rather than showing you nothing.
The evidence
Every answer is stored verbatim with the engine and the exact model that produced it, plus the sources it cited. Open any product and you can read exactly what each engine said and where it pointed, with the answering engine named on every row. Nothing is inferred or estimated — if a scan can't complete reliably, it fails loudly rather than recording a half-sampled score.
Known defects we disclose: how each engine cites its sources
Citations are not uniform across engines, and we would rather explain the difference than quietly smooth it over. ChatGPT and Perplexity return the publisher's own link. Gemini does not: its grounded answers cite Google redirect URLs (vertexaisearch.cloud.google.com/grounding-api-redirect/…) and put the real publisher in the link's title. Read naively, every Gemini citation would look like a citation from google.com — which would put Google at the top of your Top Sources and tell you to go get yourself cited on google.com.
So for redirect-style links we take the publisher from the citation's title, and that publisher is what we count, group and show you. When a redirect carries no publisher at all, it is labelled as an undisclosed source and kept out of your citation-gap list instead of being filed under Google's domain. The link itself always points where the engine actually pointed.
The labels
Aethos can sample ChatGPT, Gemini and Perplexity. Which of them is actually answering is a property of the deployment you are using, not a marketing claim: the engine strip on our homepage is generated from the running server, and Settings names the live state of every engine — answering, not switched on, no API key, daily limit reached, or erroring, with the real reason. Shipping a provider is not the same as proving one, so an engine we cannot answer with right now is never shown as live anywhere.
If an engine ever answers without citing sources, it will say exactly that rather than borrow another engine's receipts. Demo environments use clearly labeled simulated data — simulated rows carry a Simulated badge everywhere they appear, and a missing, unconfigured or budget-exhausted engine simply does not appear: it is never stood in for by a simulated answer wearing its name, and your existing scores are never overwritten to fill the gap.
What we do not measure
We do not promise that any AI engine will recommend you. We promise to measure what they actually say, show you the verbatim answer behind every number, and tell you what to fix.
And nothing here reports that an answer changed on the strength of one sample. The sampling section above states the depth every weekly figure is measured at; a claim of change is held to more than that, which is why you will not find one on any surface we ship.