Published by: MaxGrowth · GeoTrack team | Collection date: 2026-07-26 | Scope: official API web-search channels, single-round collection (see "Method" and "Scope and Limitations") Reproducibility statement: all result figures in this report come from the accompanying stats.json (regenerable by re-running analyze.py over the raw data); all question-set figures (number of questions, categories, budget phrasings) come from questions.json. The raw data is per-question JSONL containing the full answer text and the source list. We do not publish figures that cannot be independently checked.

01

Summary: Four Findings (all are observations from this sample, not general laws)

  1. In this sample, the AI named brands on almost every question. 84 of the 90 collections (93%) gave an explicit brand recommendation; among the answers that did recommend, the LLM judge extracted an average of 5.3–5.9 brands (at most 10 per answer; 2 of the 90 records hit that cap). Slots on the list are limited — whether you are on it is a competition already under way.
  2. Overlap between the three engines' recommendation lists is low. Taking as the set "the first up to 3 recommended brands to appear in the answer" (after brand-alias consolidation; not a ranking stated by the engine), pairwise: per-question mean Jaccard is only 0.25–0.39; among questions where both engines produced ≥3 brands, the sets were identical in only 0–4 questions, and had zero overlap in 6–7 questions (about a quarter). Of the 22 questions where all three engines produced ≥3 brands, only 13 had a brand common to all three. Testing a single engine does not support a conclusion about "how AI sees you."
  3. Source ecosystems are clearly stratified: in the answer citations of Doubao (ByteDance), Zhihu is at zero. In the search results of Qwen (Tongyi Qianwen, Alibaba), Zhihu-affiliated domains (Zhihu is China's Quora-like Q&A platform) account for 34.4% (99/288); in DeepSeek's retrieval substrate (Bocha, a Chinese search API), they account for 27.0% (63/233); yet across the 113 citation annotations in Doubao's answers (deduplicated by URL within each question), Zhihu does not appear once. Doubao more often cites Sina-affiliated domains (32; Sina is a major Chinese web portal) and SMZDM (23; a deals and product-review site). The same piece of content can be a core source in one engine and never surface in another.
  4. Answer formats differ by engine. Doubao and DeepSeek more often give numbered ranked lists (13/30 and 15/30); Qwen does so in only 5/30. Qwen's answers average about 800 Chinese characters, the other two about 1,500–1,600.
02

1. Method

  • Question set: 30 questions = 10 consumer categories × 3 question types (recommendation, "which brand is good" / decision, "how do I choose" / scenario, "with a budget of X yuan, what should I buy"). Categories: Bluetooth earbuds, robot vacuums, home coffee machines, dash cams, sunscreen, laptops, air fryers, running shoes, child car seats, electric toothbrushes. All question wordings are in Appendix B and questions.json.
  • Engines, channels, and the unit in which "source" is counted (real calls made on a single day, 2026-07-26; 90 calls, all successful; 🔴 "source" means three different things for the three engines — this report presents them separately throughout and makes no same-unit comparisons): | Engine | Channel | Unit in which sources are counted | |---|---|---| | Qwen (Tongyi Qianwen) | Alibaba Cloud Bailian official API, web search enabled (forced_search) | Web search results (search_results returned by the API) | | Doubao | Volcano Ark (ByteDance) official API (/responses + web_search plugin, fast-thinking tier) | Answer citation annotations (url_citation, deduplicated by URL within each question) | | DeepSeek | DeepSeek official API (deepseek-v4-flash) + Bocha search retrieval augmentation | External RAG retrieval input (the Bocha sample; not annotated by DeepSeek itself) |
  • Brand extraction: for each of the 90 answers, an LLM (temperature=0) extracted "brands explicitly listed as recommendations" (in order of appearance in the answer, at most 10 per answer), ignoring passing mentions and platform names; before cross-engine comparison, brands were consolidated through an alias table (e.g. Sony / 索尼). Extraction quality was spot-checked by script on a random sample (seed=7, 4 samples; the sample IDs and verification results are in the judge_spot_check field of stats.json): every extracted brand could be located in the original answer text (including matching Chinese and English spellings).
  • What "top 3" means: most engine answers give no explicit ranking; the three-brand set in this report is the first 3 taken by the judge in order of appearance in the answer. Pairwise comparisons include only those questions where both engines yielded ≥3 brands.
  • Persistence: each record contains the full answer text (not a summary), the source list, and the collection timestamp; a minimum-length threshold was applied to judging inputs.
03

2. Finding 1: In This Sample, a Brand List Is Written Out on Almost Every Question

Of the 90 collections, 84 gave an explicit brand recommendation (Qwen 28/30, Doubao 27/30, DeepSeek 29/30); the 6 that gave none were mostly "how do I choose" questions handled as pure spec explainers. Among the answers that did recommend, the judge extracted an average of 5.3 (Qwen) / 5.4 (Doubao) / 5.9 (DeepSeek) brands.

What this means for brands: on the consumer-category questions covered here, "will AI name brands" is no longer the open question — it names them on nearly every one. The question worth verifying is whether your brand is on the list, and how it is described; that requires question-by-question testing in your own category, and the category sample in this report cannot substitute for it.

04

3. Finding 2: Each Engine Is Writing Its Own List

On the same question, the three engines' sets of "the first up to 3 recommended brands to appear" overlap little:

Engine pair Comparable questions (both ≥3 brands) Per-question mean Jaccard Pooled (micro) Jaccard Sets identical Sets with zero overlap
Qwen ↔ Doubao 23 0.39 0.31 4 questions 6 questions
Qwen ↔ DeepSeek 26 0.26 0.23 0 questions 6 questions
Doubao ↔ DeepSeek 23 0.25 0.21 1 question 7 questions

Of the 22 questions where all three engines yielded ≥3 brands, only 13 had a "brand appearing in all three."

The implication for brands is direct: the sentence "we're doing well in AI" has to state which AI. Steady visibility in one engine's answers does not imply visibility in another; monitoring has to be done per engine, and a single-engine conclusion carried into a report will most likely be distorted.

05

4. Finding 3: Stratified Source Ecosystems — the Three Engines Lean on Different Pools

When the three engines answer the same batch of questions, the composition of their respective "sources" (units defined in the Method section; different units, presented separately) differs sharply:

Qwen (search results) Doubao (answer citation annotations, deduplicated) DeepSeek (Bocha retrieval input)
Total source records 288 113 233
Zhihu-affiliated 99 (34.4%) 0 63 (27.0%)
Sina-affiliated 51 32 6
SMZDM 21 23 24
JD.com on-site pages 12 0 4
Distinct domains 49 48 96

A few points worth noting (all of them describe composition within each engine's own column; the three columns use different source units, and their raw counts must not be compared across columns):

  • Zhihu's share within each engine's own sources varies enormously: 34.4% in Qwen's search results and 27.0% in the Bocha retrieval substrate — the largest content community in each — while it does not appear even once (0) in Doubao's answer citation annotations. In other words, the same content platform can go from "number one" to "absent" depending on the engine.
  • SMZDM is one of the few shopping-guide sites that appears in all three source sets (it can be seen in Qwen's search results, in Doubao's answer citation annotations, and in the Bocha retrieval substrate). The only claim made here is the qualitative fact that it appears in all three; the three units differ, so raw counts should not be compared to decide which is larger.
  • The distinct-domain counts (Qwen 49 / Doubao 48 / DeepSeek-Bocha 96) serve only as a within-engine structural reference, not as a cross-engine ranking of "who is more long-tail" — the metric is affected both by each engine's own source unit and by its total record count (the three columns' totals, 288/113/233, already differ), so it cannot be used to judge whose substrate is more long-tail.

What this suggests for content planning (as a hypothesis to be tested, not a proven placement effect): where you are absent from one engine, first check what source types that engine actually surfaces in your category, and add content to the corresponding pool first — whether this really changes the recommendations must be verified by re-testing the same question set before and after placement.

06

5. Finding 4: Differences in Format

  • Share of ranked-list answers: DeepSeek 15/30, Doubao 13/30, Qwen 5/30. Doubao and DeepSeek more often give an outright "first place, second place," which makes it obvious whether a brand made the top slots.
  • Source coverage (each in its own unit): Qwen 30/30, 9.6 per record; DeepSeek (Bocha input) 30/30, 7.8 per record; Doubao 26/30, 3.8 per record — 4 Doubao answers returned no citation annotations, mostly "how do I choose" explainer answers. The units differ, so no cross-engine ranking of "who is more transparent" is made here.
  • Answer length: Qwen averages about 800 Chinese characters; Doubao and DeepSeek about 1,500–1,600.
07

6. Three Action Implications for Brands

  1. Monitor per engine: break "AI visibility" into per-engine metrics, and record mentions, position, and sources separately; do not extrapolate from a single engine.
  2. Segment content by pool (a hypothesis still to be tested): prioritize content according to the source distribution each engine actually surfaces in your category; supply on the Zhihu side mainly corresponds to the Qwen / Bocha-type substrate, while for Doubao you have to look at the pool it actually annotates; the effect must be confirmed by comparative re-testing.
  3. Re-test as a routine: this report is a single-round snapshot, and AI answers fluctuate by nature; conclusions should be maintained by re-running a fixed question set periodically, recording changes — including regressions — as they are.
08

7. Scope and Limitations (Read Before the Numbers)

  1. The API surface ≠ the App front end: this issue used the official API web-search channels, which differ from the mobile App front end in retrieval strategy and answer format (product cards, for instance); for observations under the App front-end scope, see our separate hands-on study.
  2. Single-round collection: each question was collected once per engine; the figures are used to show structural differences among the three engines, not to rank brands precisely, and they cannot establish each engine's everyday general behavior.
  3. Category sample: 10 consumer categories do not represent every industry; B2B, services, and high-ticket low-frequency categories may differ.
  4. The three engines' "source" units differ (search results / answer citation annotations / external RAG input) and are presented separately throughout; DeepSeek's sources reflect the Bocha retrieval substrate, not annotations by DeepSeek itself.
  5. Brand extraction is LLM-assisted: output is in order of appearance and capped at 10 (2 records hit the cap); input hygiene and scripted spot-checking were applied (all correct), but individual brand consolidations (a sub-brand versus its parent brand, for example) may involve borderline judgments.
  6. This report does not constitute an evaluation of or a recommendation for any brand; the brand names listed are objective occurrences within the engines' answers.
09

Appendix A: Data File Manifest (archived with the report, checkable record by record)

File Contents
questions.json The 30 question wordings (category × question type; the source of the question-set figures)
results_qwen / results_doubao / results_deepseek .jsonl Per-question raw records: full answer text, source list, collection timestamp
judged.jsonl Per-question brand extraction results
stats.json All result statistics (regenerable by re-running analyze.py; the sole source of the result figures in this report)
collect.py / judge.py / analyze.py Collection, extraction, and statistics scripts (the spot-check includes sample IDs and the seed)
10

Appendix B: The 30 Questions

Bluetooth earbuds: which brand is good (budget ¥500) / how to choose noise cancelling / affordable picks for students; robot vacuums: which brand works well / how to choose (vacuum-and-mop combo) / picks under ¥3,000; home coffee machines: which type should a beginner buy / which brand of fully automatic machine is good / ¥1,000-class espresso picks; dash cams: which brand is reliable / how to choose and which specs to look at / picks under ¥500; sunscreen: which brand works well / what to use for oily skin in summer / how to choose for sensitive skin; laptops: picks for a university student with a ¥6,000 budget / which brand of thin-and-light is good / how to choose one for office work; air fryers: which brand is good / how to choose and what capacity / picks around ¥300; running shoes: which brand to start with / how to pick between cushioning and support / picks under ¥500; child car seats: which brand is safe / how to choose / picks around ¥2,000; electric toothbrushes: which brand is good / sonic or oscillating, which is better / picks around ¥200.


MaxGrowth (maxgrowth.ai) provides AI visibility monitoring and GEO (generative engine optimization) services to brands going global and to brands operating in China; GeoTrack is our visibility monitoring system. If you want to know where your brand stands in these three engines' answers, you can claim a free growth diagnostic on our site.

Want to know where your brand ranks in your own category questions?

Get your free Growth Audit