Short answer: most big AEO tools (Peec AI, Profound, Ahrefs Brand Radar, Semrush) scrape the public ChatGPT and Perplexity interfaces. Smaller tools, including AEO Copilot, call the official APIs. Scraping gets you closer to what an anonymous visitor sees. The API gets you a baseline you control, can repeat, and are allowed to collect.
The catch is that a raw API call is not a fair stand-in for the interface. The settings decide how close you get. In a 40-answer test I ran today, turning web search on changed about 4 in 10 of the brands ChatGPT recommended, and setting the user's country to the UK instead of the US changed the payroll answer almost completely.
My position: the API with the right settings is the better tracking method. Not because it copies the interface perfectly, but because nothing copies "the user" perfectly, including scraping.
Who scrapes and who uses the API
Here is what each vendor says publicly, checked on 2026-09-28.
Peec's own docs say something worth reading twice: for a logged-out user, the platform decides which model to use and whether to search. The scraper does not choose. ChatGPT does. Malte from Peec confirmed on Reddit that logged-out interface scraping is their default, and that he believes Profound works the same way.
I make AEO Copilot, so weigh my view on this accordingly. I have tried to keep the facts about other tools to what they state themselves.
What scraping captures that the API does not
Scraping has real advantages, and I don't want to pretend otherwise.
The product layer. The ChatGPT interface adds things no API returns: shopping cards, brand entity panels, the sources flyout, and the query fan-out ChatGPT runs behind the scenes. Cloro's ChatGPT scraping guide lists these and explains how they read them from the response stream. Ricardo Batista, cloro's founder, says logged-out scraping is the closest thing to what users see, that it returns the sources, and that his team never got the API to work well enough to use it.
ChatGPT's own routing. In the interface, ChatGPT picks the model and decides whether to search. Through the API, you pick. If you pick badly, you measure something no user ever sees. And the model behind ChatGPT's default answers has not always been usable with web search through the API: for GPT-5, Peec's Malte summed it up as either the wrong model, or the right model without search.
The divergence is large. Surfer compared 13,779 API and scraped answers from 1,000 prompts (collected 2026-08-04). Brand overlap between API and interface answers was only 15.5% to 23.8%, or 21.3% to 31.6% after merging brand-name variants (published 2026-09-25). For ChatGPT, the API named 13.8 brands per answer against 7.9 in the interface.
That gap showed up in my own test too.
My test: settings move the answer as much as the method
I ran 5 buyer prompts ("best CRM for a small business", "best payroll software for a small company", "best AEO tools for agencies", and 2 more) through OpenAI's Responses API on 2026-09-28. Each prompt ran 2 times under 4 settings: no web search, forced web search, forced search with a New York location, and forced search with a London location. That is 40 answers. I then extracted the recommended brands from each answer and compared the lists.
Overlap here means shared brands divided by all brands named across the 2 lists.
3 things stood out.
Search pulls the API list into the interface's range. Without search, the API named 10.1 brands per answer from training data, with no sources. Forced search brought that down to 6.4, close to the 7.9 Surfer measured in the real interface. Search is not the whole story: Surfer's ChatGPT API calls searched on 83% of prompts and still named 13.8 brands. Model choice, instructions and location all move the answer too, and in my test location moved it most.
Search made answers more stable, not less. Run-to-run overlap went from 63% to 88% with search on (71% once a location was added, method below). Answers anchored to live pages repeat themselves more than answers pulled from memory.
Location changes the answer more than anything else. Without a location, the payroll answer opened with "For most small U.S. companies" and put Gusto first. With a London location, Sage Payroll came first, followed by Xero, BrightPay and HMRC's own free tool. The US and UK payroll lists shared 8% of their brands. Across all 5 prompts, US and UK answers overlapped 48% (method). The CRM list barely moved (77%), while freelancer accounting fell to 31% as the UK answer switched to FreeAgent and Sage.
!Bar chart of how much ChatGPT API answers overlap between a New York and a London user location, by prompt: CRM 77%, AEO tools 61%, project management 61%, freelancer accounting 31%, payroll 8%. For payroll, the US top 5 is Gusto, QuickBooks Payroll, OnPay, Patriot Payroll and ADP RUN; the UK top 5 is Sage Payroll, Xero Payroll, QuickBooks Payroll, BrightPay Cloud and HMRC PAYE Tools. Only QuickBooks appears in both.
Search has a price. A forced-search call used about 22,000 input tokens on average. The same prompt without search used 16.
What is the most accurate setting?
There is no single user to copy. Every method picks a user, whether it says so or not.
!Grid comparing who picks the settings in 3 ways of collecting ChatGPT answers. Logged-out scraping: ChatGPT picks the model and decides whether to search, location comes from the proxy's IP, no memory, shopping cards visible. Logged-in scraping: ChatGPT picks within the account's plan and decides on search, location comes from the proxy and account, 1 account's history applies, shopping cards visible. Official API: you pick the model, force search on and set the location, no memory, shopping cards not returned.
Logged-out scraping shows you the average anonymous visitor, in whatever country the proxy sits in. No memory, no custom instructions, no chat history. Real ChatGPT users who are logged in have all 3, and OpenAI says memory and past chats shape answers. So scraping is not "what your customer sees" either. It is what a stranger with no history sees. Logging the scraper in doesn't fix that: 1 account's history represents nobody, and it is the riskiest option legally (more below).
The API shows you a clean baseline with no personalization, where every setting is written down. It will not match any single person. It will match itself next month, which is what tracking needs.
SparkToro and Gumshoe asked 600 volunteers to run the same prompts 2,961 times on their own devices and accounts (published 2026-01-28). The chance of getting the same brand list twice was under 1 in 100. Their conclusion was that rank is unreliable, but visibility share across many runs is a reasonable metric.
That is the frame I use. You are not trying to screenshot what 1 user saw. You are measuring how often you show up across repeated runs of a controlled setup, and whether that number moves. I wrote more about that in signals vs KPIs for AI search.
For that job, these are the settings that matter most, in order:
- Web search on. Without it you are measuring training data, which stops at the model cutoff and ignores anything you have published since.
- Location set to your market. In my test, with no location, the model assumed a US user.
- No persona in the prompt. "As a CFO at a 50-person company" is not how people type, and it changes the answer.
- Repeat runs over time. 1 run tells you almost nothing. The trend does.
- Same settings every time. Change 1 setting and your chart shows a jump that has nothing to do with your content.
Is scraping against the terms of service?
Mostly, yes, for the consumer products. I am not a lawyer and this is not legal advice.
- OpenAI's Terms of Use say you may not "automatically or programmatically extract data or Output".
- Perplexity's terms ban using any robot, spider, crawler or scraper to collect data from the service.
- Anthropic's consumer terms ban automated access, except through an Anthropic API key, and ban bypassing protective measures.
- Google's terms ban automated access that ignores machine-readable instructions like robots.txt.
The scraping side's answer is that a logged-out visitor never accepted those terms. There is some case law behind it. In Meta v. Bright Data (January 2024), a court found Meta's terms did not cover logged-out scraping of public data. In hiQ v. LinkedIn, hiQ lost on breach of contract, partly over fake logged-in accounts, and settled with a permanent injunction in 2022. And Google's case against SerpApi over bypassing its anti-bot system is still open as of September 2026.
The pattern: logged-out collection of public pages is a grey area that has held up so far. Logged-in accounts and getting around bot protection are where it goes wrong.
There is also a practical side. Scraping ChatGPT at scale means residential proxies, browser fingerprint patching and handling Cloudflare challenges. Cloro estimates (updated 2026-09-15) building it in-house at roughly $4,800 to $14,100 a month for 500,000 to 1 million requests, and notes the ChatGPT endpoint they rely on changed twice in 2025. When the interface changes, the data stops until someone fixes the scraper. When an API changes, you get a deprecation notice.
What the official API gives you
Permission. Anthropic's consumer terms ban automated access and then name the exception: an API key. That is the whole difference. No proxies, no account bans, nothing to explain when a client's legal team asks how the data was collected.
Control. When I first ran ChatGPT through the API with automatic tool choice, it often skipped the search and answered from memory. 2 runs of the same prompt could come from different places, and the chart moved for no reason. Forcing the search fixed it. That is a decision you can only make through the API. In the interface, ChatGPT makes it for you.
Reproducibility. With forced search, 2 runs of the same prompt shared 88% of their brands in my test (method). Next month I can send the exact same call and compare. You can't rerun a scraped session from last Tuesday.
Cost and speed. Scraped answers take 30 to 45 seconds each, per cloro's FAQ. My forced-search API calls averaged about 20 seconds, and a call without search about 16.
Sources. With search on, the APIs return citations: OpenAI as URL annotations, Perplexity as a citation list, Claude always, Gemini through grounding metadata. The "API has no sources" argument was true for a no-search call. It isn't true anymore. As one API-based vendor noted, all 4 major engines can search the web through their APIs. It just costs more.
What you give up: the product layer (shopping cards, entity panels), ChatGPT's own routing, and anything the consumer app adds on top of the model. If your category lives in shopping cards, scraping sees something the API can't.
What AEO Copilot does
AEO Copilot calls the official APIs for all 4 engines. Here is exactly how, as of 2026-09-28:
- ChatGPT: OpenAI's Responses API. On paid plans you can turn on live web search per account or per brand. When it is on, search is forced on every call, because with automatic tool choice the model often skipped the search and answered from memory, which made runs inconsistent.
- Claude: Anthropic's API with its web search tool when live search is on, limited to 1 search per answer. Anthropic doesn't let you force the search, so Claude decides per prompt.
- Perplexity: the Sonar API, which searches the web on every call and returns its citations.
- Google AI Overviews: Gemini, instructed to answer in the format of an AI Overview. This is the least faithful of the 4, since it doesn't reproduce Google's search results page. I would rather say that plainly than let a chart imply otherwise.
Every engine gets the same neutral instruction: answer as a helpful assistant, as a list, with numbered sources. That keeps the answers comparable across engines and easy to parse. The cost is that formatting differs from the consumer apps, which in my test used tables in all 40 answers.
AEO Copilot doesn't set a location yet, so its answers lean toward a US view. For a US brand that matches the default. For a UK or German brand, my test says it is the gap that matters most, and the one I would close first.
You can see the engine-by-engine results this produces in the 1,600-prompt study, and how the tools compare on price in AEO Copilot vs Peec vs Profound.
How to choose
If your category lives in ChatGPT's shopping cards, you need scraping, because the API can't see them. For everything else I would pick the API with fixed settings.
Whichever you pick, ask your vendor 3 questions. Is web search on? Which location do you use? Are you logged in or out? If they can't answer, the chart is measuring something, just not something they can name.
How I made this
The 40-answer test ran on 2026-09-28 through OpenAI's Responses API with low reasoning effort, no system prompt, and search forced where noted. Brand names were extracted from each answer by a second model call and normalized by parent brand (for example "Zoho CRM" and "Zoho" count as 1). 5 prompts and 2 runs per setting is a small sample: treat the percentages as direction, not precision. Vendor methods and terms were checked on the same date.