How a publishing house secures 9,378 backlinks using a self-built anchor

A travel guide contains thousands of references, all of which must remain valid in every new edition. Michael Müller Verlag in Erlangen has developed its own tool for this purpose and explains what it learnt along the way, via a defunct file format.

In the Tuscany travel guide published by Michael Müller Verlag, the text refers 9,378 times to other parts of the book – to a telephone number, a point on a map, an index entry, or a mention three chapters further on. Each of these links must survive the next revision, whether in the book, the app, the e-book or on the web. For years, this depended on a Word function that Microsoft eventually discontinued. The publisher subsequently developed its own method, calling it ‘Semantic Anchor’, and presented it at a developers’ conference in Pordenone in September. The publisher’s founder, Michael Müller, explains how an editor experiences this in their day-to-day work, which other publishers face the same problem, and where automation ends. He also explains why he does not want to entrust his authors’ texts to a public chatbot.

jumping-arrow-white

Published: 6 October 2026 | Key visual: Magnific, Screenshots: Michael-Müller-Verlag  | AI generated translation of the original article

The silent problem

They call it a ‘silent problem’: cross-references that get lost in the editing process. When did you realise that things couldn’t go on like this?

When we began developing our mmtravel app in 2008/2009, it quickly became clear that we needed a solution with a functioning ‘round-trip’ process. The tagged and published texts must subsequently be able to serve as a reliable source file for the authors of the next edition.

We had previously developed rudimentary versions for Symbian, Windows Mobile and Palm OS in collaboration with MobiPocket. This worked, but involved an enormous amount of manual work. Hotels, for example, were extracted as lists from Word documents, and custom character formats separated the hotel name, body text, telephone number and address from one another during export.

When Word 2003 offered what was thought to be an elegant solution in the form of Custom XML, we built an extensive framework of our own on top of it. However, Custom XML was anything but plug-and-play. It was precisely at this point that it became clear that we needed a more robust solution to ensure the long-term viability of our digital production.

9,000 editorial links in a single travel guide. What on earth is all linked together there?

Our IT specialist, Karsten Seidel, has broken this down for this interview using our Tuscany travel guide as an example. We have a total of 9,378 tags or links, spread across 63 chapters. The largest categories are:

  • Database POIs: 3,466

  • POIs in the document: 3,273

  • Telephone numbers: 1,989

  • Web addresses: 1,310

  • Index entries: 852

  • Headings: 562

  • Map numbers: 469

  • Images: 455

  • In-text citations: 249

For POIs, for example, we store the database ID and information indicating which entry is the main entry. This is important if a location appears more than once in the text. When creating the output formats, the system needs to know whether the link leads to the main entry, to a map in the app, or, in the e-book, simply to the central database entry.

And what happens to these links in a new edition?

We can handle a large proportion of this ourselves. Some links need to be reassigned, for example where texts have been significantly altered or sections have been moved. We now use a great many automated processes for this. However, it is not possible to do without some manual checks. With several thousand links per title, quality assurance remains an important part of the process.

Travel guides are an extreme example, but not an isolated one. For which publishers is the linking problem just as significant?

Wherever large volumes of content are updated regularly yet still need to remain properly linked. For example, in encyclopaedias, technical documentation, legal texts or software manuals.

In my view, however, travel guides are the icing on the cake. They bring together continuous text, addresses, geographical data, maps, images, cross-references, itineraries and database entries. And all of this has to be accurate in every new edition, whether in the book itself, the app, the e-book or on the web.

What the anchor looks like in the text

What does an editor see of your semantic anchor when she restructures a chapter?

In day-to-day editorial work, the technology should remain as unobtrusive as possible. It is crucial that you can see directly within the text where the semantic anchors are located and which category they belong to.

In Word 2003, we have so far been using field functions and XML structures for this. This works from a technical point of view, but it is cumbersome. The markers take up space and can even alter the line breaks when they are shown or hidden.

01_Screenshot_Interview_dpr_Semantic_Anchor

This is what the anchors look like in Word 2003. Image: Michael Müller Verlag

In LibreOffice, this is handled much more elegantly. The anchors appear as narrow square brackets, can be hidden and do not affect line breaks. Depending on their category, they have different colour shades, making it easy to distinguish between POIs, references and other tags. Technically, everything is contained within the OpenDocument file itself. The anchors are stored as bookmarks within the text and are managed in a separate RDF file within the same container.

03_Screenshot_Interview_dpr_Semantic_Anchor

The same heading in LibreOffice. Image: Michael Müller Verlag

When a chapter is being revised, the anchors remain visible and can be managed directly within the relevant content. At present, it is mainly our tagging team that is using this feature. In future, authors, editors and layout designers will also be able to make direct use of these functions.

08_Screenshot_Interview_dpr_Semantic_Anchor

The costly lesson

Did anything go wrong during development?

Our original system was based on Word 2003’s custom XML tags. Microsoft subsequently discontinued the feature we were using. This meant we were dependent on a technology that was central to our production but was increasingly becoming obsolete.

For years, we have tried to keep this system running, using VMware ThinApp, specialised Office installations, Microsoft’s MSIX technology, CrossOver on the Mac and, at times, Office for Mac. Technically, we were able to salvage a great deal, but the effort involved became ever greater. Today, such old versions of Office bring with them additional problems relating to security, updates and compatibility. Our tools are therefore effectively ‘deprecated’. With open interfaces and standardised formats, we are significantly more independent when new technical solutions need to be found.

So what happened next?

Fortunately, we had László Németh on the team, a developer who laid the technical groundwork for a user-friendly display of the anchors. Building on this, Karsten Seidel and Rose Haberecht developed a comprehensive extension in Python. It was a rocky road, partly because we lacked the convenience of a familiar development environment whilst debugging. But anything is possible.

However, this new approach is not entirely without its limitations. For further production, our documents are also processed in InDesign. This involves exporting them to Word format, which results in the loss of the semantic tags. However, the unique IDs can still be reused, so the links are not completely broken.

Advertisement

Newsletter-Banner_Xpublisher_2000x694

The chatbot: the next instalment

Your press release explicitly mentions chatbots as a delivery channel. What role does AI play in this concept?

We have already simulated this feature in Claude, and it worked well. However, in order to use it with our authors’ texts, it needs to be made to run in conjunction with an isolated LLM, so that the authors’ texts do not become public domain.

The Federal Ministry of Research, Technology and Space has classified the method as eligible for funding. Will it be made freely available, and what does a publisher need if it wishes to use it?

LibreOffice and the extensions we have developed using the LibreOffice source code are open source. Building on this, we have developed our own proprietary plug-in that allows texts to be tagged and structured in a targeted manner. The underlying open-source platform remains fully intact.

Anyone wishing to familiarise themselves with the technology will soon be able to do so with no obligation. We will be making a publicly accessible community version available, which will allow users to try out the key features and test their own use cases. We then intend to use this as a basis for developing solutions for professional use cases.

You have been in the publishing business since 1979 and have witnessed a number of technological changes in the industry. What sets this one apart from the previous ones?

From the ball-head typewriter to the laser printer, via the mmtravel app, and now to AI: what sets this transformation apart is, above all, its pace. In the past, individual work processes were digitised or new means of production were introduced. AI has the potential to transform a great many areas simultaneously, from research and editing to planning, personalisation and delivery.

The expectation is that, all of a sudden, everything is supposed to work in the blink of an eye. I find it incredibly exciting to see whether that will actually happen and where this development will take us. In any case, I’m very curious to see where the wind will take us this time.

What publishers can take away from this

Michael Müller’s experience is particularly relevant to publishers who reprint their content several times.

  • Dependencies include: Which function, on which production depends, is available from only one manufacturer? At Michael Müller Verlag, it was a single Word function, and its removal meant it took years to find a workaround.

  • Planning the return journey: Published data must be incorporated into the next edition as a clean source. Anyone who only creates versions for the app and e-book has to rework every edition by hand.

  • Treat links as part of the inventory: 9,378 references in a single title constitute a dataset with its own quality assurance process. Automated processes help, but the final checks are still carried out manually.

A short video demonstrates how Semantic Anchor works in practice Demo video from the publisher. Michael Müller would be delighted to receive any questions, feedback or ideas for further areas of application in the comments section.

MM_vor_dem_Pauli-Brunnen_auf_dem_Erlanger_Marktplatz_Ausschnitt_2 (1)

Michael Müller (LinkedIn) founded the Michael Müller Verlag in Erlangen. Today, the publisher is one of the leading providers in the German-language travel guide sector and has a catalogue of around 220 titles.

From an early stage, Müller focused on combining traditional travel content with digital applications. The first version of a digital travel guide was created as early as 2007; with the mmtravel app, the publisher developed this into a fully-fledged service that combines editorial content with interactive maps, tours and other digital features.

The article is part of the Channels: Digital Publishing Technologies, which focuses on content strategies and processes. The channel is sponsored by Fabasoft Xpublisher.