One might expect language models to promote their own brands due to corporate instructions or training data. However, a practical test of five systems with enabled web search showed unexpected results: AI assistants are reluctant to put themselves on a pedestal and frequently recommend competitors' products.
The Core of the Experiment
The test involved five platforms connected to real-time search engine results: ChatGPT (gpt-4o), Claude (claude-sonnet-4.5), Gemini (gemini-3.6-flash), Perplexity (sonar-pro), and Alisa AI (Yandex Search generative answer).
The services were given four neutral prompts without mentioning specific brands (for example, about choosing an assistant for writing or solving complex tasks), as well as one control question directly listing the participants. Each query was repeated three times to account for response variability. Out of 75 queries sent, 71 returned valid results.
How Models Evaluate Themselves
In blind tests (without hints in the query), participants displayed varying degrees of "modesty":
- Claude mentioned itself in all 12 cases, but placed itself in first position only three times;
- ChatGPT mentioned itself in 9 out of 12 answers (4 times in first place);
- Gemini named itself 10 out of 11 times, but never put its name at the top of the list (median position was third);
- Perplexity mentioned itself in 6 out of 12 answers (3 times in first place);
- Alisa AI named itself in only 3 out of 10 successful runs.
ChatGPT Is the Common Favorite
An analysis of all 57 valid responses to neutral questions revealed that AI models actively recommend third-party solutions. ChatGPT became the leader in citations — it was included in recommendations even by direct competitors (Claude, Perplexity, and Gemini mentioned OpenAI's product in nearly 100% of cases).
| AI / Brand | Occurrences in 57 answers |
|---|---|
| ChatGPT | 51 |
| Claude | 47 |
| Gemini | 42 |
| DeepSeek | 30 |
| Perplexity | 23 |
| GigaChat | 15 |
| Alisa AI | 13 |
| Grok | 9 |
Systems such as DeepSeek, GigaChat, and Grok did not take part in direct testing, yet were actively mentioned by other models based on web search data.
Why This Happens
Such a result does not necessarily signify the objective superiority of one model over another. Rather, generative responses reflect the distribution of information across the internet and current search results: the more frequently a service is covered in articles and reviews, the higher the probability that a language model will cite it in a response to a user.
Comments
to leave a comment.
No comments yet.