Ask ChatGPT and Gemini the same shopping question twice, and the product named first may change. A Sept. 24 analysis from marketing firm Suff Digital found ChatGPT changed its top pick in 56.3% of back-to-back answers to identical “What is the best…?” questions, compared with 82.7% for Gemini.
Other studies produce a less clear-cut result. Recent research has measured top-pick stability, shortlist churn, factual contradictions, and source overlap, with several tests relying on API configurations rather than the consumer apps people use directly.
ChatGPT changes its top pick less often
The Sept. 24 Suff Digital analysis asked 300 “What is the best…?” questions 10 times each on ChatGPT and Gemini, producing 6,000 answers. ChatGPT gave the same top pick in all 10 runs for 8.3% of questions; Gemini did so for 1.3%.
Suff Digital says ChatGPT used web search and the queries were run from the US, but its methodology provides less detail about Gemini’s retrieval setup. The results represent one controlled comparison rather than a universal ranking of the two services.
A Product.ai study published Sept. 22 tested 220 shopping questions five times each across free and paid configurations of ChatGPT, Gemini, Claude, and Perplexity. Of the 217 questions with complete answer sets, 86% produced at least one confirmed factual conflict.
Gemini contradicted one of its own earlier answers on 29% of questions in its free configuration and 27% in paid, compared with 22% and 18% for ChatGPT. Product.ai also found at least one “costly error” on 56% of questions for Gemini free and 54% for paid, versus 19% and 17% for ChatGPT. The tests used provider APIs rather than the consumer ChatGPT and Gemini interfaces.
The reliability of AI shopping advice is becoming more consequential as assistants move closer to purchases. Apple, for example, is testing a conversational shopping assistant inside the Apple Store app that can answer product, trade-in, payment, and order questions.
Both tools still show heavy recommendation churn
MentionBird’s Aug. 27 study reached a different result after normalizing for list length. Across 38,374 ChatGPT answers and 38,830 Gemini answers to 585 commercial-intent questions, a brand moved into or out of the top five on 39.2% of ChatGPT transitions and 39.7% of Gemini transitions.
MentionBird used GPT-5 mini and Gemini 2.5 Flash through APIs rather than ChatGPT.com or gemini.google.com. On that normalized shortlist measure, the two systems were effectively tied.
A Sept. 16 academic preprint found another source of variation. Across 1,536 product-related responses, the ChatGPT and Gemini consumer interfaces shared only 5.4% of displayed source domains on average and had no domain in common in 76.7% of comparisons.
The information behind those answers can shift over time. August tracking found ChatGPT citations to Reddit fell sharply as its retrieval behavior changed, showing how search changes can reshape the sources feeding an answer.
ChatGPT was more stable in Suff Digital’s top-pick test and contradicted itself less often in Product.ai’s API study. MentionBird, however, found almost no difference in normalized shortlist churn. None of the studies supports treating either assistant’s first recommendation as settled.
IT and procurement teams can use either tool to build an initial product or vendor shortlist, but prices, specifications, model generations, and availability still need verification against vendor documentation or retailer listings. Research into chatbot warning labels also suggests generic cautions may not reliably prompt users to verify AI-generated claims independently.
Want to learn more AI tips, tricks, and prompting techniques? Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy, our practical learning platform designed to help professionals use AI more confidently at work.
Learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. Check out all the lessons here →