“Hey AI, which book can you recommend?”
Study on the sources of AI-generated reading recommendations
More and more people are asking ChatGPT or Gemini this question and are promptly given a handful of titles, each with a brief explanation. A study by the Munich-based digital agency Hattenberger Partner shows where the AI gets its reading recommendations from, which publishers are featured, and what publishers can do to be included.

Published: 29 September 2026 | Key visual: AI-generated, Magnific | AI generated article of this original article
Anyone looking for a Christmas present in 2026 – perhaps a good book – is increasingly likely to ask ChatGPT or Gemini. According to a study by McKinsey, 44 per cent of users now cite AI search as their preferred source of information, whilst traditional online search accounts for 31 per cent. But what does the AI actually recommend, and what sources does it draw on? The digital agency Hattenberger Partner wanted to find out exactly and, in August and September, asked the two AI assistants 228 questions for reading recommendations, spread across twelve categories ranging from crime novels to financial guides.
The good news is that publishers themselves constitute the largest source group. Their websites account for 1,591 citations, which is 26.4 per cent of all sources. The press and broadcast media follow with 16.2 per cent, bookshops with 13.7 per cent, and reading portals and blogs with 12.2 per cent.
This is remarkable in that a publisher’s website is, for the most part, a form of self-promotion. Nevertheless, the AI treats it as an independent source. This is particularly evident in the case of best-of lists. 216 citations refer to rankings published by publishers on their own websites, such as ‘the ten best crime novels’. Such lists contain exclusively titles from the publishers’ own imprints. The most frequently cited ranking page in the study is one of these; it ranks at rowohlt.de and totals 124 quotations.
This pays off for the publishers. If the AI cites a publisher’s website, that publisher’s books are almost always included. For Carlsen, this figure stands at 91.3 per cent; in responses that do not cite the publisher’s own website, it is 12.3 per cent. The study itself points out that cause and effect cannot be separated here. In practice, it amounts to the same thing: when the AI reads a publisher’s website, it also recommends that publisher’s books.
ChatGPT reads the arts section, whilst Gemini reads the sales charts
ChatGPT and Gemini do not take the same approach. Both rely on rankings with roughly the same frequency: for ChatGPT, one in four sources refers to such a list, whilst for Gemini it is almost one in four. However, these are different lists:
ChatGPT draws on best-of lists compiled by editorial teams, such as the SWR best-of list, the crime fiction best-of list, or the book recommendations from Deutschlandfunk Kultur and NDR. Such lists appear 277 times as evidence in OpenAI’s language model, whereas in Google’s Gemini they appear only 10 times. Gemini, on the other hand, relies on bestseller lists such as those published by Der Spiegel and on rankings from bookshops. Overall, almost one in four sources in Gemini comes from the retail sector, compared with just one in fourteen for ChatGPT.
For publishers, this means there are two ways to get the answers. A title makes it onto a ‘best-of’ list through reviews and press coverage, whilst a title makes it onto a bestseller list through sales and retail visibility. However, the recommendations vary greatly: the study asked each question on three separate occasions, and only 8.5 per cent of the titles mentioned appeared in all three sets of answers.
Biggest surprise: Amazon is missing as a source
Amazon is Germany’s largest bookseller. Among the 6,036 references used by ChatGPT and Gemini to back up their book recommendations, there are amazon.de not even once.
The declaration is contained in the robots.txt file, which a website uses to specify which programmes are permitted to crawl it. According to the study, Amazon blocks the search programmes of both AI providers in this file. Hugendubel blocks neither of them and is cited 140 times, whilst Thalia is cited 177 times. The group’s subsidiary, Audible, maintains its own file, allows the programmes access and is cited 117 times.
A second example shows that robots.txt controls more than one might think. Penguin Random House’s books top the citation statistics for several of the segments in question; Heyne, for instance, leads in the fantasy and science fiction categories. The group’s publishing imprints are all hosted at the same address penguin.de, but this address does not appear in a single citation. A look at their robots.txt file provides at least a partial answer: the entry ‘Google-Extended’ is blocked there. According to Google’s documentation, this entry governs two things at once. It determines whether content is included in the training of future Gemini models, and whether Gemini is permitted to use it as a source when answering a question. It has no effect on normal Google Search.
Anyone who wishes to exclude their content from the training will also be excluded from Gemini’s responses. Incidentally, the file does not explain the findings to ChatGPT; its search programmes are permitted to penguin.de read. “We have not been able to identify any external technical reason as to why ChatGPT still does not cite the page,” says the study’s author, Dr David Hanisch of Hattenberger Partner.
What publishers can do now
The study draws three recommendations from the findings, ranked in order of effort required.
Keep your own leaderboards. 216 of the study’s citations come from rankings published by publishers on their own websites. Such a page requires very little: a clear heading stating the genre and year, a numbered list, and for each title, two sentences of context in the body text, including the author’s name and year of publication. Pages consisting solely of cover images are of little help; the systems read text.
Make your own website the source. According to the study, a page is considered citable if it answers a question, such as the recommended reading age for a book or the order in which the volumes in a series should be read. The author, year, series and volume should be included in the visible text and not just in the metadata. Each publishing brand should have its own URL. Publishing previews, in the form of publicly available PDFs, also appear in the records.
Appearing where one’s own genre is referenced. The sources the AI draws on depend heavily on the genre. For non-fiction, 40.7 per cent of the references come from publishers’ websites; for novels, the media lead the way at 31.9 per cent; for fantasy, it is reading portals; and for financial guides, it is specialist portals outside the book industry. When the AI cites arts and culture sections, it is the press office that decides; when it cites reading portals, it is the community.
The study will be published to coincide with the Frankfurt Book Fair at https://hattenbergerpartner.de/
About the method
AI Visibility Study: The Book Market 2026 (Hattenberger Partner). 228 queries for book recommendations across twelve segments, posed to ChatGPT (GPT-5.6) and Gemini 3.7 Flash whilst located in Germany, both with and without web searches, on three dates between 21 August and 2 September 2026. 2,736 responses analysed, 6,036 source references from 870 addresses. Titles are assigned to the publisher of the German first edition via the DNB catalogue. Cited from the preliminary version dated 21 September 2026.
The article is part of the Channels: Digital Publishing Technologies, which focuses on content strategies and processes. The channel is sponsored by Fabasoft Xpublisher.
