Publishing without text: Frontiers turns datasets into academic articles and uses AI as an editor
...However, only a handful of these publications have actually been published in the last twelve months. Why is this, and what lessons can every publisher currently integrating AI into their production processes learn from this?

Published: 24 September 2026 | Key visual: AI-generated, Magnific
A Swiss scientific publisher has launched a product that contains not a single written sentence. Frontiers, one of the world’s largest open-access publishers, has been selling raw research data as fully-fledged publications since 2025. An AI system organises the dataset and describes it; experts review it; and a portal publishes it. The end result is a publication that can be cited. This costs around 5,500 Swiss francs, even though the same data is available free of charge in open archives. Nevertheless, only a handful of these publications have appeared in the past twelve months. The Australian library and information specialist Lilly Hoi Sze Ho reviewed the system for the trade magazine Katina. Their findings are also of value to publishers outside the academic sector.
In academia, the same hierarchy applies as in most specialist publishing houses.
The text is published; the supporting material is an appendix. Data sets, tables and raw data are deposited free of charge in open repositories, such as Zenodo, which is run by the CERN research centre, or Figshare. There, they are assigned a permanent identifier, the DOI, and remain accessible. Revenue is generated one level above that, from the article itself.
Frontiers is turning this order on its head – and making no secret of it. “90 % of Science Is Lost” was the headline the publisher used for the press release announcing the launch of its FAIR² platform, pronounced “Fair Squared”.
According to Hos’s description, the service consists of three components. AI-powered curation converts raw data into a machine-readable package and enriches the metadata. It presents each dataset via its own portal, complete with visualisations, exploration tools, a downloadable Jupyter notebook and a generated podcast. Furthermore, a peer-reviewed data article is published in a Frontiers journal, making the dataset citable.
The platform’s name alludes to the FAIR principles, according to which research data should be findable, accessible, interoperable and reusable. The superscripted square suggests that all of this should be taken to the nth degree. In practical terms, this means one thing above all: the dataset is accompanied by a ‘package leaflet’ that can be read not only by humans but also by machines.
- Advertisement -

The editor’s name is Clara; the readers are machines
At the heart of the process is an AI that Frontiers refers to as a ‘data steward’ and has named Clara.
Clara cleans and validates data, creates data dictionaries, suggests metadata fields and variable descriptions, organises the methodology documentation and checks compliance with the in-house standard. A chat window remains visible in the portal to answer questions about the schema and article structure. The metadata is exported in Croissant format, a standard developed by the MLCommons initiative that describes datasets for training and evaluation by AI models.
This makes it clear who this publication is intended for. The reader of a Data Article is, first and foremost, a model seeking training data, or a script performing a meta-analysis across hundreds of datasets. Frontiers is developing a publishing product for a machine reader, and human peer review ensures that the machine is not presented with anything incorrect. In line with this, authors must disclose where AI-generated content has been incorporated into the data or documentation.
In her article, Ho precisely outlines the limits of automation. How well Clara generates metadata depends on the quality of the original documentation, on oversight by the researchers, and on specialist knowledge that automated tools may not be able to capture. The AI describes what it finds. It does not fill in the gaps in a poorly documented dataset.
What the 5,500 Swiss francs are paid for
Storage space is free on Zenodo, whereas on Frontiers it costs 5,500 Swiss francs.
The set-up fee, equivalent to around 5,800 euros or 7,100 US dollars, applies to datasets of up to 50 gigabytes and covers AI curation, peer review, the portal and long-term hosting. It is only payable once the dataset has been accepted.
Added to this is the publication fee charged by the respective journal. When Ho wrote her review in May 2026, Frontiers did not itemise this fee separately, which is why she described the transparency of pricing as only partial and noted that the total cost might be a deterrent for small teams and individual researchers.
Since then, the publisher has revised its fee schedule. It sets out four categories ranging from 990 to 3,150 Swiss francs, graded according to the journal’s maturity and the financial strength of the subject area, specifies the fee on every page of the journal and describes what it covers. Those unable to pay may apply for a fee waiver. According to Frontiers, it had waived over nine million dollars by 2025.
More interesting than the amount of the fees is the logic behind them. What is being sold is the certification. A DOI from the repository confirms that the data exists. A peer-reviewed data article confirms that someone with specialist knowledge has examined it and that its structure meets a published standard.
In return, there is the opportunity to be cited, recognition within the academic community and proof for funding bodies, which have long since made data management plans mandatory. According to Ho, the fee can also be paid from their funding pots.
A handful of articles in twelve months
The first peer-reviewed data article was published in March 2025.
When Ho reviewed the system in February 2026, the pilot data sets had been processed, along with one article from December and two from January, covering topics ranging from coronavirus variants to biodiversity surveys in the Indo-Pacific. Just a handful, after a year of operation. Ho therefore suspects that the assessment and certification processes are progressing only slowly.
Apparently, the bottleneck lies precisely where AI is of no help. AI speeds up the processing stage, but Frontiers organises the peer-review process in the same way as it does for its text-based journals. A dataset from marine research requires different specialists and different metadata standards to one from drug discovery. According to Ho, it remains to be seen whether subject-specific support can be combined with a uniform technical architecture. It is precisely these differences that have already held back previous FAIR infrastructures.
Her further observations read like an acceptance report familiar to any publisher that has rolled out a platform.
Accessibility: The portal does not offer a translation or read-aloud function.
Interface: The API provides metadata and download links, but hardly any filtered queries.
Documentation: The guidance on Data Articles remains limited.
Registration: The ORCID login responded to several attempts with a simple error code; the workaround involved registering via email.
Dependency: Curation, publication and the portal are all managed by a single provider. This raises questions about long-term availability and integration with institutional repositories.
What does this mean for publishers?
Even if this does not concern academic publications, it is worth taking a look at the process.
There are three lessons to be learnt.
New products: Frontiers has turned a type of content that was previously regarded as a by-product into a product in its own right, with its own workflow, its own quality standards and its own price. Anyone sitting on structured content – from standard parts tables and recipe databases to course materials – has similar by-products at their disposal. The question is which of these, with certification and machine-readability, can be developed into a product in its own right.
New ‘readers’: Croissant metadata, open licences, persistent identifiers, formal specifications, disclosure requirements for AI components. None of this is designed for humans. Anyone wishing to licence content to AI systems in future will have to deliver it in the same way that Frontiers delivers its datasets. The metadata becomes the product; the content is attached to it.
New responsibilities: Ho observes this within her own profession. Because the data is openly accessible following publication, libraries no longer need to manage it, and their role is shifting towards providing advice on data management plans, publication budgets and long-term archiving. In publishing, the situation is the reverse. The more AI takes over the sorting and description, the more the publisher must take on the role of oversight. The publisher’s name stands as a guarantee that someone has checked the data.
Certification cannot be automated
Ho recommends monitoring developments and keeping an eye on dissemination, technology and sustainability.
The platform could be of value to institutions with complex datasets that are seeking visibility, re-use and recognition.
Another insight is important when it comes to publishing decisions. Frontiers has deployed AI where it encounters the least resistance: in the packaging process. However, peer review and verification must still be carried out by humans, and the pace of work in these areas remains different.
Anyone drawing up their own AI roadmap should bear this order in mind. The machine speeds up the processing stage. The review stage cannot be automated to the same extent, but it is the most important reason why you need a publisher.
German adaptation of the review ‘A New Publishing Infrastructure Treats Datasets as Formal Research Outputs. Can It Work Across Disciplines?’ by Lilly Hoi Sze Ho, published on 27 May 2026 in Katina Magazine (Annual Reviews), DOI 10.1146/katina-052726-1, CC BY-NC 4.0. With the kind permission of the author and Katina Magazine.

Lilly Hoi Sze Ho (LinkedIn) is an experienced information science specialist based in the Northern Territory, Australia. She has extensive experience in library technologies, collection management, digital infrastructure, metadata standards and technical services. Her professional interests include digital collections, open science, open infrastructure, library systems and the development of standards in library and information science. Internationally, she is involved with the International Federation of Library Associations and Institutions (IFLA), including the Advisory Board for Open Science and Scholarship and the committees of the Asia-Oceania Regional Division. She also contributes to standards and metadata initiatives. Ho publishes academic articles and works as a book and peer reviewer. Her work focuses on sustainable and interoperable approaches to library technologies, digital access and scholarly information services.
Photo of the author: private.
The article is part of the Channels: Digital Publishing Technologies, which focuses on content strategies and processes. The channel is sponsored by Fabasoft Xpublisher.
