App check: the app vs the API
Paste what ChatGPT, Gemini, Perplexity, Claude or Google AI Mode showed you, and see how far Citeroot's API answers are from the app, for your own questions.
Citeroot collects answers through each provider's official API, the same way every run. That's a consistent measuring stick, but it isn't exactly what a person sees in the app: the app adds its own instructions, can use a different model, and may use the person's history and location. Why API answers and app answers differ explains the causes.
App check measures that gap for your questions. You ask one of your questions in the app yourself, paste the answer into Citeroot, and Citeroot compares it with its own answers to the same question.
#Doing a check
- Open App check in the app and pick a question (or one of this week's suggested checks).
- Ask it in the app exactly as written, in a new chat. Signed out, or in a private window, is the fairest comparison.
- Copy the whole answer, ideally with the app's copy button so links come through, and paste it. Paste the sources the app showed too, one per line.
- Tell Citeroot the app, whether you were signed in, whether it searched the web, the country you were in and the date.
Citeroot checks the pasted answer with the same rules it uses for collected answers (brand named, recommended, your site cited, competitors named) and shows it side by side with its own answers.
#Suggested checks
Each week App check suggests a few of your active questions, each paired with an app, picked at random. People tend to check the questions where they expect a difference, which makes the gap look bigger than it is. Checks on suggested questions avoid that, and the summary says how many of your checks were suggested. A suggestion stays the same all week; new ones appear every Monday (UTC).
#How a check is paired with Citeroot's answers
A check is compared with Citeroot's answers that are:
- from the paired engine: ChatGPT with the OpenAI API, Claude with the Anthropic API, Gemini with the Gemini API, Perplexity with the Perplexity API, and Google AI Mode with the AI Mode search page;
- to the exact same question. If you reworded it, the check is shown but not counted;
- for the same country you were in. For ChatGPT, Claude and Google AI Mode the country is sent to the provider as a location hint. The Gemini and Perplexity APIs don't take one, so there it's the project's market. Answers collected before Citeroot recorded locations can't be matched to a country, and the page says so;
- collected within 14 days of your check, up to the 20 nearest in time;
- from one model. If Citeroot's answers came from more than one model, the most common one is used and the others are set aside and counted.
Simulated (sample) answers are never compared with a real app.
#The numbers
Everything is calculated per question first, then averaged across questions, so checking one question ten times doesn't count as ten questions.
Named you: app vs Citeroot. For each question, the share of your checks that named you and the share of Citeroot's paired answers that named you. The summary shows both averages and the difference in points with a 95% interval. From 10 questions the interval is a paired t-interval across questions. Below 10, Citeroot uses whichever is wider of that and Newcombe's interval for a difference of two proportions, because small samples deserve caution.
Brand difference beyond run-to-run variation. Brand overlap is the share of tracked brands named in either answer that are named in both. Answers vary from run to run, so Citeroot compares two things: how much the app's answers overlap with Citeroot's, and how much answers overlap within a channel. The extra difference is the gap between them, in points:
- When at least 5 questions were checked twice or more in the app, run-to-run variation is measured on both sides (Citeroot's repeated answers and your repeated checks). This is the sounder baseline.
- Otherwise it's measured from Citeroot's repeated answers only, assuming the app varies about as much. That baseline is weaker, and the page tells you which one was used. Asking the same question twice, in fresh chats, is the most useful thing you can do to sharpen the comparison.
Same brand named first and sites in common are calculated the same way. Sites are only compared when both sides list sources, and subdomains count as the same site.
Both sides are re-checked with your current brand and competitor names, so a competitor you added later is looked for in every answer alike.
#How to read the summary
Each row is one app, signed in or signed out. Rows are never added together, and there's no "overall" figure.
| Questions compared | Shown as | What you see |
|---|---|---|
| 1–4 | Anecdote | Counts only, no percentages or intervals |
| 5–14 | Early read | Rates, differences and 95% intervals |
| 15 or more | Estimate | The same, with narrower intervals |
The verdicts come from the interval, not the point estimate:
- More / less often only when the whole interval for "named you" is above or below zero.
- Similar rates only when the whole interval is within ±10 points.
- No extra difference in brands only when the whole interval is within ±5 points.
- Can't tell yet otherwise. That's a common and honest answer with few questions.
A row is marked provisional while any of its checks is less than 14 days old, because Citeroot may still collect answers inside that check's window. Every row lists which checks weren't counted and why, the date range, and how many checks were on suggested questions.
Copy summary for a report gives a paragraph you can paste into a client report. It includes the period, the number of checks and questions, the verdict and a link back to this page, so anyone can see how it was worked out.
#What App check can and can't tell you
- It compares the app with Citeroot on the questions you checked, around the dates you checked them. It says nothing about questions you haven't checked.
- If you mostly check questions where you expect a difference, the gap will look bigger than it is across your whole panel. Suggested questions avoid this.
- One person's app can differ from another's. Signed-in answers can depend on your history and settings, so they're shown separately from signed-out ones.
- Both sides are read with the same rules, but pasted text can include or leave out things the app shows as cards or links. Some differences come from copying, not the engine.
- Only brands you track are counted. A brand the app names that you don't track doesn't show up here.
- The app may use a different model from the one Citeroot calls. That's part of the difference being measured, and it changes over time.
- It shows that answers differ, not why.
#Your data
- Checks are stored in the project until you delete one, all of them, the project or your account. They're included in Settings → Download my data.
- What you paste is never sent to an AI provider. It's analysed on Citeroot's servers with the same text rules as collected answers.
- Share links are stored as text so you can find the chat again. Citeroot never opens them.
- Email addresses in a pasted answer are partly hidden before saving (for example
[hidden]@example.com). Remove anything else personal before pasting, especially from a signed-in session. - Research totals (optional). You can let Citeroot count a check in research on how app and API answers differ. Nothing has been published yet; when it is, it will be totals across many accounts only, never your questions, brands or answer text. It's off unless you tick it, and Stop using my checks for research on the App check page opts every check in the project out.
App check doesn't use your monthly answer allowance. Each plan includes a number of checks per month: Scout 20, Trail 200, Summit 1,000 and Expedition 5,000.
#Why not collect the apps automatically?
Some tools drive the consumer apps with automated browsers. The apps' own terms restrict that: OpenAI's terms bar extracting output "automatically or programmatically", Anthropic's consumer terms bar access "through automated or non-human means" except with an API key, and Perplexity's terms bar automated collection. Citeroot uses the official APIs for collection and puts what a real person saw next to them, so you can measure the difference instead of guessing.
Last updated Oct 9, 2026 · Suggest an edit