Your toothpaste knows what you will buy next
Colgate found a way to replace 9,300 survey respondents with AI. The market research industry has not noticed yet.
Here is something strange. A toothpaste company quietly published a paper last October that should have sent the entire market research industry into a panic. Eight months later, almost nobody has noticed.
Colgate-Palmolive teamed up with a research group called PyMC Labs and submitted a paper to arXiv. The title was dry. “LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings.” The kind of title that makes you keep scrolling. But inside that paper is a finding that, if it holds up, changes how companies decide what to sell you.
The problem they set out to solve is one that anyone who has tried to use AI for market research has bounced off of. Ask a large language model to rate a product on a scale of 1 to 5. It gives you back a 3. Ask it again. Another 3. Ask it a hundred times and every answer clusters around the middle. The responses are flat. They are safe. They are useless. This has been a wall. Researchers have been throwing themselves at it for years.
The Colgate team found a way around it that is so simple you wonder why nobody thought of it sooner.
Here is how it works. You do not ask the AI for a number. That is the trap. Instead, you give it a person. A 34-year-old woman in Chicago, household income 85,000, two kids, shops at Target. You show her a new product concept. Then you ask her to write down what she thinks. Raw. Unfiltered. As that person.
The AI writes a paragraph. “I like that it does not have a strong smell. I tried something similar last year and it gave me a rash, so I am cautious. But if the price is right I would give it a shot.”
That paragraph gets converted into a vector. A long string of numbers that represents the meaning of the text. The vector gets compared against anchor statements. A response that reads like “I would absolutely buy this, this is exactly what I need” lands closer to the 5 anchor than the 1 anchor. The rating emerges from semantic distance. Not from asking the AI to produce a number it cannot honestly compute.
They called this Semantic Similarity Rating. SSR.
Then they tested it. The dataset was Colgate’s own research pipeline. 57 product surveys. 9,300 real human responses, collected the old fashioned way, with panels and incentives and field time. They ran the same surveys through SSR and compared the results.
The synthetic consumers matched real buying behavior at 90% of human test-retest reliability. The distribution of their responses was statistically almost indistinguishable from the human panel. A Kolmogorov-Smirnov similarity above 0.85. Translation: if you showed a market researcher the two sets of numbers, they could not tell you which came from real people and which came from the AI.
The synthetic consumers did something else that the researchers did not expect. They gave better qualitative feedback than the real humans. Real survey respondents rush through questions. They want the five dollar gift card. They write one word answers. The AI, unbothered by time or compensation, wrote detailed, critical, sometimes harsh evaluations. The synthetic consumer in Chicago did not like the packaging. The synthetic consumer in London thought the price point was off.
VentureBeat called this “the dawn of the digital focus group.”
Think about what this means for an industry.
The market research business is roughly 80 billion dollars. It is built on panels of real humans. Recruiting them. Screening them. Paying them. Waiting for them. A typical concept test takes weeks to field and costs thousands of dollars per demographic segment.
SSR collapses that to an API call and a batch job that runs overnight.
You want to test pricing sensitivity across six income brackets. Run it. You want to know how 25-year-old men in Texas feel about a product. Run it. You want a thousand interviews with single mothers in the Pacific Northwest. Run it. The bottleneck shifts from finding people to asking smarter questions.
This is where the story gets complicated, which is where it gets interesting.
Synthetic consumers are not real consumers. They are statistical ghosts trained on internet text. They reproduce known patterns well. They will fail on genuinely novel products. The Colgate dataset was all personal care. Toothpaste. Deodorant. Things people have strong familiar opinions about. How well does SSR work on something nobody has seen before. A new category. A technology that does not exist yet. The paper does not answer that question.
Ninety percent accuracy also means ten percent error. For a new flavor of toothpaste, that margin is fine. For a drug launch or a billion dollar product line, it is not. The method that generates clean synthetic data can also generate convincing but wrong answers. The difference is hard to spot because the wrong answers look exactly like the right ones.
There is a deeper problem that keeps me up. Homogenization. If every company uses the same LLMs to simulate consumers, do products start to converge toward the same average preference. Stanford GSB published a paper in 2024 showing that AI generated survey responses tend to be suspiciously nice. They lack edge. They lack the snark of real human feedback. The synthetic consumer that always gives reasonable feedback might produce a world of reasonable but indistinguishable products.
PyMC Labs published their own write-up on the methodology, framing SSR as a response to the Stanford problem. If uncontrolled AI pollutes human datasets, controlled AI can generate clean ones from scratch. Defense becomes offense.
The honest take: market research does not die tomorrow. SSR becomes the new first pass. You screen concepts with synthetic panels, fast and cheap. Then you validate the survivors with real humans. The cost of the wrong product going to market drops because the cost of discovering it is wrong drops first. That alone is a revolution.
The paper is still on arXiv. Open access. Anyone can read it. Anyone can build on it. Colgate published the methodology that could eat their own research budget. That is corporate generosity or the most honest signal yet that the way we make decisions about what to sell is about to change.
A toothpaste company just showed the 80 billion dollar market research industry how the next decade works. The swarm is here. Whether it replaces the focus group or just makes it cheaper to find the right one depends on how honestly companies use it.
And here is the strangest part. Eight months after the paper dropped, the industry is still running focus groups. Still paying panels. Still waiting weeks for data. Nobody has panicked yet. You have to wonder how long that lasts.


