
If you’re trusting Google’s Gemini to track down online deals or verify product specs, you might want to pause before hitting that buy button. A new benchmark test from the data team at Product.ai revealed that both free and paid tiers of Google Gemini are currently the most likely among leading AI chatbots to make up product claims and deliver costly mistakes to shoppers.
The team put popular chatbots through 220 straightforward e-commerce questions across nine categories. Then they repeated every single question five times to check for consistent answers. The lineup included Gemini 3.6 Flash and Gemini 3.1 Pro Preview, OpenAI’s GPT-5.6 Luna and GPT-5.6 Sol, Anthropic’s Claude Sonnet 5 and Claude Opus 5, along with Perplexity‘s agent API across subscription tiers.
Invented ingredients, wrong prices, and flip-flopping answers
Testing the paid tier with gemini-3.1-pro-preview uncovered surprising reliability issues. The model hallucinated a fake product claim—like an invented ingredient, wrong spec, or bogus model name—21% of the time, leading all major platforms in fabrications.
Even worse for deal hunters, 56% of Gemini’s responses contained at least one mistake that would cost a buyer real money. This included instances of quoting incorrect price points or pointing to the wrong item.
| Engine · tier | Questions with at least one likely-fabricated claim | Questions with at least one costly error |
| Gemini · free | 20% | 56% |
| Gemini · paid | 21% | 54% |
| Claude · free | 22% | 44% |
| Claude · paid | 12% | 21% |
| ChatGPT · free | 6% | 19% |
| ChatGPT · paid | 7% | 17% |
| Perplexity · free | 4% | 15% |
| Perplexity · paid | 3% | 14% |
To make matters stranger, Gemini frequently contradicted itself, regularly giving entirely different answers when asked the exact same product question just five minutes later.
| Engine · tier | Contradicted its own prior answer, at least once |
| Gemini · free | 29% of questions |
| Gemini · paid | 27% of questions |
| Claude · free | 26% of questions |
| ChatGPT · free | 22% of questions |
| ChatGPT · paid | 18% of questions |
| Claude · paid | 17% of questions |
| Perplexity · free | 17% of questions |
| Perplexity · paid | 13% of questions |
Product.ai data expert Dakota Nunley pointed out that since these language models operate on probabilities, their inability to give consistent answers minutes apart shows how risky it is to treat them as automated shopping advisors.
Perplexity takes top marks while old-school shopping stays king
By contrast, Perplexity pulled off much better accuracy numbers across both free and paid setups. Its paid agent API tier only made up product details 3% of the time, while keeping financial pricing errors down to 14%. It also proved to be the most consistent model when answering repeated queries.
At the end of the day, the researchers suggest using AI for what it handles best: brainstorming ideas, discovering brands, and vetting purchase concepts. However, smart consumers should still double-check details directly on the manufacturer’s site or shop around the old-fashioned way when exact prices, ingredients or model numbers matter.
The post Shopping with AI? Benchmark Tests Find Gemini Is Most Likely to Mislead Buyers appeared first on Android Headlines.
​Â