Key Takeaways
- Published estimates of consumer AI use in 2026 range from 42% to 85%. The studies aren't measuring the same behaviour, which is why the range is so wide and why no single figure can be trusted as a planning baseline.
- Self-reported media use is one of the least reliable measures in behavioural research, and AI use is a harder recall task than anything before it: high-frequency, low-salience, fragmented across tools that blur together, and subject to social desirability pressure.
- When we asked participants to declare their AI use and then watched six days of screen recordings, every participant was running more than one assistant, and every participant got something about their own week wrong.
- AI has introduced a third consideration state alongside considered and rejected: never surfaced. A brand can be absent from a decision without ever being evaluated, and no analytics tool can see it.
- The practical fix isn't a better survey question. It's observation: capture the session, treat the declared-versus-observed gap as a finding, and log the whole answer set including the absences.
Why don't AI adoption statistics agree?
Because they're measuring different things. Ask about a month or a year, about shopping or about any use at all, about using AI somewhere in a journey or about replacing search entirely, and you get four different numbers from four honest studies. The range in the 2026 record runs from 42% to 85%, and the spread is a property of the questions, not of the consumers.
Here's the simple version of the question. What proportion of consumers use AI when they shop?
The research record answers it like this. NIQ puts it at 42% having used at least one AI tool to shop in the past month. Alchemer reports 48.5% having used an AI tool to research a purchase in the past year. Stord says 51%. Capgemini finds 58% have replaced traditional search engines with generative AI for product and service recommendations. McKinsey's ConsumerWise survey has 68% of US consumers using at least one AI tool in the past three months. And a SEMrush survey of 1,030 US shoppers reports 85% using AI at least weekly, though that one comes with a caveat we'll come back to.
The obvious reading is that somebody is wrong. That's not the interesting reading, and it isn't the one we'd defend.
Look closely and these studies aren't measuring the same thing. One asks about a month, another about a year. One asks about shopping, another about any use at all. One asks about replacing search, which is a much stronger claim than using AI somewhere in a journey. Of course they produce different numbers.
Two rows are worth pausing on. The 85% is the figure most often quoted as evidence that AI shopping is near-universal, and it isn't measuring that. SEMrush screened for people who had already tried AI and excluded the quarter of the sample who hadn't, so the base is AI users rather than shoppers, and the 85% covers weekly use of AI for anything at all. The figure for weekly product research is 55%. Both are honest numbers. Only one is an adoption rate, and it isn't the one that travels.
The 68% points the other way, and the caveat is McKinsey's own. They note the figure is likely an underestimate, because people use AI-based search and AI features inside apps without registering it as AI use. One of the six studies is telling you in its own footnotes that its instrument can't see the whole behaviour.
That's the problem. Two years into the largest change in consumer information behaviour since mobile, the industry hasn't settled on what the behaviour is, let alone how much of it is happening. And there's a reason it can't settle: almost every one of those figures comes from asking people. On this particular behaviour, asking doesn't work.
Why can't people accurately report their own AI use?
Because AI use is close to the worst-case scenario for human recall. It happens many times a day, it leaves almost no memorable trace, it's spread across tools that feel interchangeable, and there's mild social pressure in both directions. People will answer the question. The answer won't be reliable.
None of this is new or controversial. It's one of the better-established results in communication research. Parry and colleagues, publishing in Nature Human Behaviour in 2021, ran a pre-registered meta-analysis across 106 effect sizes comparing logged digital media use against what the same people said about their own use. Self-report correlated only moderately with the logs, and rarely reflected them accurately. Their conclusion was blunt: findings that rest on self-reported use alone should be treated with caution.
Scharkow's earlier validation study, comparing survey responses against client-side log data from a large household sample, found the same divergence and added a detail that matters here. The misreporting was systematic rather than random, and under-reporting and over-reporting had different predictors. They aren't one error with one direction. They're two errors with two mechanisms. Revilla, Ochoa and Loewe found the same pattern in metered web behaviour, where some destinations were systematically under-reported and others systematically over-reported.
This is the part that should worry anyone building a strategy on survey base numbers. If the error were a constant, you could calibrate it. Take the survey number, apply a known correction, move on. But an error that runs in both directions for different reasons can't be corrected out. You don't have a number that's too high or too low. You have a number of unknown sign.
So what makes AI use a harder recall task than the social media use that literature was built on?
It's high-frequency. Frequent behaviours are the ones people estimate worst, because there are too many instances to count and the mind substitutes a general impression for an actual tally.
It's low-salience. Opening a chat window to check something isn't an event. It's a utility, closer to reaching for a light switch than to watching a film. There's no episode encoded, so there's nothing to retrieve.
It's fragmented, and this is the new one. Social media use was spread across apps with strong distinct identities. AI use is spread across tools that do broadly the same thing through broadly the same interface, and they blur.
There's a fourth mechanism working alongside the memory problem. A 2026 CHI paper on the underreporting of AI use argues that self-reports are shaped in part by social desirability bias, the tendency to answer in a way that feels acceptable rather than accurate. Sometimes that pushes people to overstate how fluent they are with AI. At work, it more often pushes them to understate how much they lean on it. Either way it's a second distortion sitting on top of a recall task that was already failing.
Put those together and you're asking someone to summarise something they never encoded, spread across tools they can't distinguish in memory, with a mild incentive to shade the answer.
The clearest demonstration of how far that gap can open comes from an unrelated field. METR ran a randomised controlled trial in 2025 in which 16 experienced developers completed 246 real coding tasks, half with AI tools permitted and half without. Before starting, they predicted AI would make them 24% faster. Afterwards, having done the work, they estimated it had made them 20% faster. Measured against the clock, they had taken 19% longer, a result the researchers verified partly by analysing screen recordings of the sessions. That was a snapshot of early-2025 tools and METR has since published newer data, so it says nothing durable about AI and productivity. What it does establish is narrower and more useful: competent professionals were confidently, measurably wrong about their own recent experience, and it took observation to show it.

What happens when you watch people use AI instead of asking?
You find that people are running several assistants rather than one, and that they can't accurately describe their own week. When we showed participants their own screen recordings, they corrected themselves without prompting.
Earlier this year we ran a study called The Real Feed. One arm of it, an AI deep dive with seven participants across the US and Spain, was built specifically to test the question above. Small sample, and we're not presenting it as a population estimate. What it can do is show the mechanism working.
The Real Feed, AI deep dive: study design. Day one, participants told us which AI tools they use and what they use them for. Then six days of screen recordings capturing actual sessions. Day six, they watched themselves back and told us what they'd got wrong. The design was deliberately awkward for participants, and that was the point.
Two things came out of it.
The first was structural. Every participant was running more than one assistant. Not one person used a single tool. The alias each participant filed under ended up reading as an inventory: GPT, Gemini, Copilot. GPT, Perplexity, Gemini, Claude.
The second was the audit, and it's the finding. Participants didn't describe their own week accurately, and when shown the evidence, they said so themselves.
One participant named ChatGPT on day one, then corrected himself: he hadn't been precise, and Copilot is what he actually uses for everything. A second discovered she had leaned on Gemini far more heavily than she'd claimed, and said plainly that her day-one answer was probably wrong. A third summarised her week by saying she'd thought she used AI less than she does, and in fact uses it constantly. A fourth found her usage wasn't spread evenly across her tools the way she had assumed. She was reaching for one far more than the others without having noticed.
None of these people were being careless, and none had a reason to mislead us. They were doing their competent best at a question that isn't answerable from memory. That makes it a measurement problem, not a participant problem. It's the same dynamic as our screen recording research on digital influence, where the story people tell about what influenced them is far tidier than the trail their screens record. With AI the behaviour is more frequent and less memorable, so the gap opens wider.
Is there a single AI journey to map?
No. Participants route different questions to different tools, and the routing rule sits below the level of conscious report. There's no one journey to map, there's a set of them, allocated by a logic the person can't articulate.
The stack finding deserves its own beat, because it quietly invalidates a lot of current planning.
The industry says "AI" and pictures one destination. Often it says "ChatGPT" and means the whole category. What our participants actually did was route: quick factual questions to one tool, multi-step research to another, a second opinion to a third when the first answer felt thin. One described going broad first and then following whichever branch interested him, deliberately avoiding long specified prompts because he wanted to choose his own path through the answer.
And nobody could describe their own routing accurately. The routing is the behaviour. It determines which brands a person encounters, and it sits below the level of conscious report.
Which means "our AI journey" is a category error before the research starts.
What is the third consideration state AI has created?
Never surfaced. Research has always tracked two outcomes: a brand was considered, or a brand was rejected. AI adds a third state in which a brand is simply absent from the answer, never evaluated and never dismissed, and nothing in your analytics can see it.
The strongest evidence we have for this wasn't collected on purpose.
Screen to Shelf was a pet food shopping study: eleven participants across the UK and Germany, twenty-six tasks covering the cupboard, the last purchase, the online research, the shopping trip, the debrief. Not one task mentioned AI.
AI turned up anyway.
Shoppers described AI Overviews as increasingly useful and as the first thing that catches their attention on a results page. One said the AI answer always draws her eye and that she likes seeing what it agrees with, an unusually clear description of a system being used as a confirmation device rather than an information source. Another found AI-generated review summaries handy for getting the gist of what customers were saying. A third was sharper: she pointed out that when most reviewers say a food doesn't smell and a few say it does, the summary tends to present that as an evenly weighted problem, simply because of how it words things.
And brands moved through it. Purina, Lily's Kitchen, Naturo, Harrington's: entering, confirming or dropping out of consideration inside an AI answer, before a single brand or retailer site was opened.
Three states, not two. Considered: the brand was evaluated and stayed in. Rejected: the brand was evaluated and ruled out. Never surfaced: the brand was absent from the answer set, so it was never evaluated at all. The first two leave a trace. The third doesn't.
A participant in the AI study spelled it out. Asked what a brand would learn from watching his week, he said his usual protein brand would probably be disappointed.
Across the entire conversation about protein, it never came up once.
Not evaluated and dismissed. Simply absent.
L.E.K. Consulting put a number on the consequence: 31% of AI users say their purchase decision was largely made before they reached a brand or retailer site, up from 26% two years earlier. If a third of decisions are substantially settled upstream of anything you own, the state that costs you most is the one your analytics can't see. There's no bounce to measure on a page that was never loaded. For teams already mapping the omnichannel path to purchase, this is a new upstream stage that has to be observed rather than inferred.
Trying to see how AI sits in your customers' decisions? Talk to our team about running a screen recording study. If you don't have the capacity in-house, our Catalyst team can handle design, recruitment, moderation and analysis. Talk to our team.
How should you research AI use instead?
Stop asking and start watching. Four changes cover most of it: capture the session rather than the tool list, treat the gap between declared and observed behaviour as a finding, look for AI inside the studies you're already running, and log the whole answer set including the brands that never appeared.
Screen to Shelf is the clearest argument for the third row. It found more about AI in the purchase journey than a dedicated AI study would have, precisely because nobody prompted for it. The same logic applies to routine and diary studies more broadly: the behaviours worth understanding tend to show up when you're not asking about them directly.
The second row is easier to act on than it sounds. A day-one task asking what people think they do, a period of screen recording, and a final task where they review their own footage. If you already run diary studies, you can add the audit step without redesigning anything.

Are researchers already treating AI as a behaviour?
Yes, and the shift is visible in our own project base. Across projects run on the platform between January 2024 and March 2026, work that studies AI as a consumer behaviour has gone from under 2% of 2024 projects, to under 3% of 2025, to 8% in the first quarter of 2026. Roughly five times its 2024 share, and the direction has been consistent enough that we no longer read it as noise.
The composition matters more than the growth. Almost all of that work treats AI as the behaviour under investigation, not as a way of doing the investigating. Banks studying how people use generative AI for money decisions. Retailers studying AI-assisted shopping across apparel, home and garden, and hard goods. A carmaker studying in-car voice agents. A publisher testing an AI audio prototype. A twelve-market study of AI in creative workflows.
There's a loud industry argument about whether AI can stand in for research participants. It's a real argument. It's also, on this evidence, not the one our clients are spending their budgets on. They're researching AI because their customers are using it.
One more signal, and it's the methodological tell. Screen recording appears in around 40% of those AI studies, against about 6% of everything else on the platform. Nobody mandated that. Researchers arrived at it because it's the only place the behaviour is visible. It's the same instinct behind mobile ethnography generally: if you want to understand what people do, get closer to the moment they do it.
Which is the whole argument, in the end. The question isn't whether to research AI. Your customers have already decided that for you. The question is whether the instrument you're using can see what they're doing - and if it depends on them remembering, it can't.
Frequently asked questions
Why do AI adoption statistics vary so much?
Because the studies measure different constructs, over different reference periods, on different base populations. One asks about AI use in the past month, another about the past year, another about replacing search engines entirely. Published figures range from 42% to 85%, but the 85% is weekly use of AI for any purpose among people already screened as AI users, which isn't an adoption rate at all. Before comparing two numbers, check what each one asked, over what window, and who was in the sample.
Can you measure AI use accurately with a survey?
Not on its own. Self-reported digital media use has been shown to correlate only moderately with logged behaviour, and the error runs in both directions for different reasons, so it can't be corrected with a fixed adjustment. AI use is a harder recall task than social media use because it's frequent, unmemorable, and spread across tools that blur together in memory. Surveys are useful for attitudes and stated intent. For behaviour, they need observation alongside them.
How do you research AI in the purchase journey?
Screen recording is the most direct method, because the behaviour happens on a screen and leaves no other trace. Ask participants to record actual sessions rather than describe them, and capture the full answer set: which brands appeared, in what order, with what framing, and which never appeared at all. It's also worth looking for AI inside broader shopper studies rather than commissioning a dedicated AI project, since unprompted AI use tends to be more honest than prompted AI use.
What does "never surfaced" mean in AI-assisted shopping?
It's a third consideration state. Traditionally a brand was either considered or rejected, and both leave a measurable trace. When an AI assistant answers a shopping question, a brand can be absent from the answer entirely, so it's never evaluated and never rejected. There's no site visit to analyse and no bounce to measure, which makes it invisible to standard analytics and visible only through observation.

