Structured survey participants using real, everyday accounts.
The NZ AI Answers Project
What AI really tells
New Zealanders
Submitted AI answers collected and cleaned to an explicit denominator.
AI engines: ChatGPT, Google AI Mode and Microsoft Copilot.
Fieldwork window, 2026. Anonymised, aggregated and reproducible.
Search Everywhere Optimisation is the new SEO
76 participants used real accounts to ask the same commercial questions of ChatGPT, Google AI Mode and Microsoft Copilot, then repeated the prompts in the ChatGPT Temporary condition.
They returned 888 submitted answers. The engines split on the lead provider in 134 of 211 matched cases. The same person received at least one different ChatGPT lead in 50 of 72 comparable answers.
Everything published here is anonymised, aggregated and tied to an explicit denominator.
Co-Founder of SparkToro and Moz
“I think this is wonderful work. I’m thrilled to see other people running these analyses.
Your findings speak to a fear I’ve had since January, which is that AI mention tracking, even when done at scale, with the right number of runs and metrics, fails to capture the complexity of the statistical lottery at play, including personalization and non-recommended mentions.
I fear many marketers and execs are being misled by their AI tracking tools. Hopefully studies like yours will help.
At least a few folks wake up.”
One clean run is a benchmark, not the whole picture
Unique structured survey participants
Submitted AI answers
Cross-engine splits: 134 of 211 matched cases
People with ≥1 lead change: 50 of 72
Accounting leads naming Xero
AI visibility is a distribution across engines, people, account conditions, prompts and time. A single clean run is a useful benchmark, not a complete description of what customers received.
01
What AI told New Zealanders
The top answer doesn’t reflect real market share — and it changes depending on which AI engine you ask.
The engine changes the winner
Same person, same prompt: AI engines produced different lead providers depending on the category.
The split varies by category. Power and KiwiSaver frequently produced different lead providers across AI engines, while Account remained consistent.
*Cross-engine comparisons use matched cases only: the same participant, answering the same category, with an analysable lead from all three engines. This produced 71 matched KiwiSaver cases, 70 power cases and 70 accounting cases: 211 in total.
KiwiSaver has three engines and three leaders
*Why the sample sizes differ: results are calculated from available, category-correct responses that named an identifiable first provider. All 76 participants produced an analysable ChatGPT KiwiSaver response. Google AI Mode had one response that named no provider ( =75). Copilot had four unavailable responses and one response that named no provider (n=71).
Power winners split sharply by engine
*Power denominators differ by engine after cleaning. ChatGPT had 76 analysable responses. Google AI Mode had one wrong-category paste and two responses with no identifiable lead (n=73). Copilot had four unavailable responses (n=72).
Accounting converges on Xero
This category behaved unlike KiwiSaver and Power as every output named the same provider.
*Accounting denominators reflect available, category-correct answers. ChatGPT had 76 analysable responses. Google AI Mode had two wrong-category pastes (n=74). Copilot had four unavailable responses and one wrong-category paste (n=71). Every one of the resulting 221 analysable leads named Xero first.
Presence and lead position tell different stories
The big four banks appear but never lead
Visibility without salience: ASB appeared in 97 outputs, yet the lead-position distribution is held by Kernel, Milford, Simplicity and InvestNow.
AI leans slightly away from the big four parent groups
Ownership changes the interpretation
Retired brands persist in AI answers
Normal engines alone: 77 of 221 answers mention Frank Energy (34.8%).
Account condition changes the answer
50 of 72 people received at least one different lead provider
61 of 215 paired category runs changed the lead provider
184 of 215 paired category runs changed the wider brand roster
Temporary Chat removes access to saved memories for personalisation. Other account, custom-instruction, system and run-to-run effects may still operate, so the report treats this as a condition comparison.
*Normal-Temporary comparisons include only cases where the same participant produced an analysable ChatGPT lead in both conditions for the same category. This yielded 72 KiwiSaver pairs, 72 power pairs and 71 accounting pairs: 215 pairs in total.
The wider shortlist changes more often than the lead
Accounting keeps Xero in front while three quarters of paired answers change the brands surrounding it.
*Normal-Temporary comparisons include only cases where the same participant produced an analysable ChatGPT lead in both conditions for the same category. This yielded 72 KiwiSaver pairs, 72 power pairs and 71 accounting pairs: 215 pairs in total.
A small source ecosystem recurs across power answers
Source recurrence is an association. The study does not identify which source caused an answer.
Single-run benchmarks did not always match the human plurality
Human plurality from real-account normal-engine answers versus one Peec and one Scrunch run per engine/category.
| Category / engine | Human plurality | Peec | Scrunch |
|---|---|---|---|
| KiwiSaver – ChatGPT | Kernel – 56/76 | Kernel | Milford |
| KiwiSaver – Google | Milford – 68/75 | Milford | Milford |
| KiwiSaver – Copilot | Simplicity – 39/71 | Simplicity | Milford |
| Power – ChatGPT | Ecotricity – 35/76 | No provider | Ecotricity |
| Power – Google | Powershop – 31/73 | Energy Online | Electric Kiwi |
| Power – Copilot | Ecotricity – 59/72 | Ecotricity | Ecotricity |
| Accounting – ChatGPT | Xero – 76/76 | Xero | Xero |
| Accounting – Google | Xero – 74/74 | Xero | Xero |
| Accounting – Copilot | Xero – 71/71 | Xero | Xero |
*Denominator (n) = available, category-correct responses containing an identifiable first-named provider. Unavailable responses, wrong-category pastes and answers that named no provider are excluded. The applicable n is shown for every engine and comparison.
The headline rates stabilised before the final respondent
Cumulative structured survey sample in collection order; valid denominators vary at each checkpoint;
From n=50 onward, both cumulative headline rates remained within one percentage point of their final values.
What this means for businesses
The measurement design follows directly from the observed distribution.
Measure leads and mentions
A brand can appear widely without occupying the lead position. Report both, with counts and the correct denominator.
Plan by engine
KiwiSaver and power produced different champions on ChatGPT, Google AI Mode and Copilot.
Separate brand from owner
A subsidiary can dominate the answer while its parent master brand barely appears in front.
Audit the recurring source set
Canstar, Powerswitch, Consumer NZ and a short list of specialist sources recur across answer text.
Build rebrands into answer-layer work
Frank’s persistence shows that commercial retirement and answerlayer retirement can move on different timelines.
Use clean tracking as a benchmark
A platform run is one controlled layer. Add repeated real-account testing when the customer experience matters.
What this means for consumers
An AI answer is one observed outcome, not a market verdict.
Ask another engine
For KiwiSaver and power, a second engine often produces a different first-named provider.
Verify with independent sources
Use comparison services, regulators and provider information before making a financial or utility decision.
Check the provider still exists
The answer layer can retain brands after their customer proposition changes.
Treat confidence as presentation
A polished answer can hide the variation visible across engines, people, conditions and time.
The study measures recommendation salience. It does not decide which provider is best for an individual.
02
How the study worked
Four measurement layers, fixed prompts, explicit exclusions and a denominator for every claim.
Four layers measure four different conditions
The study combines experienced visibility with controlled benchmark layers.
These layers complement one another; none is a complete population estimate on its own.
The protocol held the questions fixed
The study combines experienced visibility with controlled benchmark layers.
Start a new chat
Use the exact wording. No follow-ups. Paste the complete answer into the collection form.
Run three normal engines
ChatGPT, Google AI Mode and Microsoft Copilot using the contributor’s everyday access condition.
Repeat in Temporary Chat
Run the same three prompts in ChatGPT Temporary Chat. Twelve potential answers per structured respondent.
Every result carries its own denominator
The answer universe narrows as availability, category fit, and analysability rules are applied.
| Universe | Count |
|---|---|
| Unique participants | 76 |
| Designed human cells | 912 |
| Unavailable positions | 24 |
| Usable human outputs | 888 |
| Wrong-category pastes | 5 |
| Category-appropriate outputs | 883 |
| Analysable lead positions | 879 |
| Matched three-engine cases | 211 |
| Normal–Temporary pairs | 215 |
| Clean-layer outputs | 18 |
*Denominator (n) = available, category-correct responses containing an identifiable firs -named provider. Unavailable responses, wrong-category pastes and answers that named no provider are excluded. The applicable n is shown for every engine and comparison.
Limits define the claim, not the value of the study
The design measures reproducible salience across a specific set of prompts, accounts, engines and dates.
1. Convenience sample
Recruitment was network- and referral-based. The study does not estimate prevalence for all New Zealanders.
2. Three fixed prompts
Prompt wording defines the commercial slice. Other tasks and phrasings can produce different distributions.
3. One 14-day window
Models, sources and interfaces change. The tool layers are single-run snapshots.
4. Lead is a proxy
First-named provider is a salience proxy, not recommendation strength or suitability.
1. Repeat the field wave
Track whether winners and disagreement rates move as engines update.
2. Broaden prompt families
Test task, phrasing and journey-stage sensitivity within each category.
3. Recruit stratified samples
Separate account, subscription, occupation and demographic conditions with sufficient power.
4. Audit source pathways
Use controlled tests before claiming publisher causality.
Frequently asked questions
What is the NZ AI Answers Project?
The NZ AI Answers Project is a July 2026 field study by The Optimisers in which 76 New Zealanders used their real accounts to ask ChatGPT, Google AI Mode and Microsoft Copilot the same commercial questions, producing 888 answers analysed to an explicit denominator.
Do different AI engines recommend different brands?
Yes. For the same person and prompt, the three engines named a different lead provider in 63.5% of matched cases (134 of 211), with power and KiwiSaver varying most and accounting staying consistent.
Which KiwiSaver provider do AI engines recommend in New Zealand?
Which power company do AI engines recommend in New Zealand?
The lead flips by engine — ChatGPT (46%) and Copilot (82%) most often named Ecotricity, while Google AI Mode most often named Powershop (43%).
Which accounting software do AI engines recommend for NZ small businesses?
All three engines converged on Xero — every one of the 221 analysable accounting answers named Xero first.
Does appearing in an AI answer mean a brand is recommended?
No. Presence and lead position differ: Electric Kiwi was mentioned in 97% of relevant outputs but was the first-named provider only 25% of the time, and the big four banks appeared but never led (0 of 222).
Does your account or chat mode change the AI's recommendation?
Yes. Comparing each person’s normal account with ChatGPT Temporary Chat, 69% received at least one different lead provider and 86% saw the wider brand roster change.
Does your account or chat mode change the AI's recommendation?
Yes. Comparing each person’s normal account with ChatGPT Temporary Chat, 69% received at least one different lead provider and 86% saw the wider brand roster change.
Are single-run AI visibility trackers accurate?
They are useful benchmarks but not complete. Against the human plurality, one Peec run matched 7 of 9 engine/category cells and one Scrunch run matched 6 of 9, so The Optimisers recommend adding repeated real-account testing.
Get the full 28-page report
We’ll email you the complete NZ AI Answers Project PDF — every finding, the method and the references in one document.
We’ll email the PDF and occasional AI-search insights.
References & cite-this-study
Primary evidence and selected context used in the published report.
The Optimisers (2026). The NZ AI Answers Project: what AI really tells New Zealanders. Auckland, August 2026. theoptimisers.com
