An AI citation study is only as good as its methodology, so we're publishing ours before we collect anything. This is The Frontier Search Report, our quarterly study of which pages AI systems cite, and this page is the pre-registration: 120 commercial queries, four AI surfaces, every cited source recorded, plus a control group of pages that rank for the same queries but never get cited. The queries are frozen. The analyses are frozen. When the results land, they'll land here, on this URL, next to the promises we made before we saw the data.
That ordering is the whole point. Read on for why, and for exactly what we're testing.
A note on the name
This study was pre-registered on 2026-08-01 as the Narsil AI Citation Study. On 2026-08-03, before any data collection started, we renamed the series The Frontier Search Report, because it's the research arm of frontier SEO: the practice of optimizing for the newest AI search systems by measuring their behavior instead of guessing at it.
To be precise about what changed: the name. The research question, the 120 frozen queries, the four surfaces, the coding scheme, and the five committed analyses are untouched, and the frozen methodology file records the rename in a dated addendum. A rename after seeing data would deserve your suspicion. A rename before collection is cosmetic, and we're disclosing it anyway, because that's the standard we're asking you to hold the rest of the industry to.
Why does AEO research need a pre-registration?
Because most of it is marketing wearing a lab coat.
The typical GEO study works backwards: an agency notices a pattern in a handful of answers, writes "AI loves schema" or "AI loves long content," and ships the post with no query list, no sample size, and no way to check the work. The next study finds the opposite. Both get cited in sales decks.
Two failures repeat across almost all of it:
- No control group: finding that cited pages "answer questions directly" is meaningless if the uncited pages do too. Without comparing against pages that rank but don't get cited, you're describing the internet, not citations.
- Post-hoc queries: if you pick the queries after seeing results, you can produce any finding you want. Frozen query sets remove that lever.
Pre-registration fixes both by making the promises public before the data exists. If our findings end up boring, we publish boring findings. That's the deal.
What exactly are we testing?
One question: when AI surfaces answer commercial questions, which pages get cited, and what separates them from pages that rank for the same query but don't?
The setup, in numbers:
| Element | Spec |
|---|---|
| Queries | 120, frozen before collection |
| Industries | Dentists, law firms, home services, restaurants, real estate, B2B SaaS, plus cross-industry consumer questions |
| Intent types | Recommendation, urgent near-me, cost, comparison, trust |
| Surfaces | Google AI Overviews, ChatGPT with search, Perplexity, Microsoft Copilot |
| Locations | 8 named US metros of varied size, plus national queries |
| Control group | Google top-10 pages not cited by any surface |
| Collection window | All queries, all surfaces, within 7 days |
The industries aren't random. They're the buying situations where AI recommendations decide real revenue: someone asking for a dentist, a lawyer, a plumber, a restaurant, an agent, or a piece of software. No query mentions our brand, our clients, or any competitor. We want to see the answers as they are, not our reflection in them.
How will the collection work?
Logged-out, clean-profile sessions, US location, one 7-day window. For every query on every surface we record each cited source in order of appearance, screenshot the answer, and separately log Google's top 10 organic results. A surface that returns no AI answer gets recorded too; the no-answer rate is a finding, not a failure.
Then every unique page, cited or control, gets coded blind on the same features: page type, whether it answers the question in the first three sentences, schema present, word-count band, geographic scope, visible freshness, and whether it's owned by a recommended brand. Blind means the coder sorts pages alphabetically and codes without knowing which ones were cited. It's tedious. It's also the difference between data and vibes.
What analyses did we commit to?
Five, frozen now:
- Citation concentration. Which domains and page types soak up the citations, per surface and pooled.
- Cited versus control. Feature rates among cited pages against ranked-but-uncited pages. No claim survives into the writeup without passing this comparison.
- Cross-surface overlap. Whether ChatGPT, Perplexity, Copilot, and AI Overviews cite the same pages or live in different worlds.
- Rank dependence. What share of citations come from the query's Google top 10, which tests the common claim that AI answers are built from search results.
- No-answer rates. How often each surface declines to answer commercial queries at all.
We also committed to what we won't claim: no causality, no precision theater around a 120-query sample, and nothing about industries we didn't test. Directional findings, stated as directional.
What do we expect to find?
Straight answer: we have hypotheses and they might be wrong. Based on the mechanisms we've written about, query fan-out and grounding behavior, we'd guess citations skew toward listicles and comparison pages, overlap heavily with organic rankings, and favor pages that answer fast. We're publishing those guesses here so you can watch them survive or die. A pre-registration where the authors already know the answer is just a press release with footnotes.
If the data says schema doesn't matter, we'll say so, even though we deploy schema for clients. If it says rankings barely predict citations, that reshapes our own method, and we'll say that too.
When do results land?
After the wave-1 collection window runs, and we haven't started it yet: all 120 queries get collected within one 7-day window, and we'd rather tell you that plainly than invent a date. When they land, wave 1 results will replace this section, on this URL, with the pre-registration text preserved below them for accountability. Summary tables ship with the writeup. The full query list is already frozen, and The Frontier Search Report re-runs quarterly with the same queries so changes over time are measurable instead of anecdotal. Same name, same queries, every wave.
In the meantime, the aggregated side of the Frontier Search Report is already live: our AI search statistics page maintains the field's key numbers with every source linked and dated, and the wave-1 findings will be summarized there alongside the full writeup here.
Until then, two things you can do. If you want the findings the moment they publish, the answer arrives fastest by checking back or asking us directly. And if you want to know where your business stands in AI answers before our data arrives, run the test yourself tonight: ask ChatGPT, Perplexity, and Google the questions your customers ask, and write down who gets named. That baseline will make the results worth more to you when they land.