LIVE
+++ Commentary: AI firms grab training data at any cost +++ EdgeEver puts an MCP server inside your note archive +++ PiClaw puts the Pi coding agent in a browser workspace +++ LLMs as a Cognitive Virus: Preprint Models Dependence +++ AIHawk: an open-source browser agent that speaks MCP +++ OpenGUI: an Android agent that taps through real apps ++++++ Commentary: AI firms grab training data at any cost +++ EdgeEver puts an MCP server inside your note archive +++ PiClaw puts the Pi coding agent in a browser workspace +++ LLMs as a Cognitive Virus: Preprint Models Dependence +++ AIHawk: an open-source browser agent that speaks MCP +++ OpenGUI: an Android agent that taps through real apps +++
All news ›
AI IN LIFE AI IN LIFENEWS
DAILY
DE EN
MODELS

Commentary: AI firms grab training data at any cost

A heise opinion piece tallies the tactics: crawlers that ignore robots.txt, books bought only to be scanned and binned, medical talks mined.

Commentary: AI firms grab training data at any cost

Symbolic image: a robotic arm turns the page of an open book under a camera rig, trimmed book spines piling up beside it in front of a server rack with blinking status lights.

In a September 5, 2026 opinion column for heise online, Wolf Hosbach argues that AI vendors now cross every line to obtain training data, from ignored robots.txt rules to recorded conversations with doctors.

At a glance

  • Opinion column by Wolf Hosbach, published September 5, 2026 in the commentary section of heise online.
  • Central charge: the collectors' servers push past the rules site operators publish in robots.txt.
  • Coding agents comb through Git projects unasked and can end up executing malicious code, the piece says.
  • Doctolib wants to use content from recorded doctor consultations for AI training, per reporting by Zeit.
  • Publisher Carlsen is suing OpenAI over the style of Astrid Henn's Neinhörner books, the column reports.

Writing on September 5, 2026, heise online columnist Wolf Hosbach argues that the companies training large language models have stopped recognizing any limit on where their data comes from. This is an opinion column, not a news report, and its examples are assembled to carry that thesis.

Why the appetite keeps growing

Large language models need very large volumes of original human text, and material recycled out of other models only goes so far. That demand is the column's starting point. When a lab needs fresh text, it takes text from wherever text happens to sit.

A robots.txt file is a request, not a lock

The first charge concerns ordinary web crawling: collection servers walk past the instructions operators leave in robots.txt. The file was never an access control, and it works only for as long as everyone chooses to honor it. The column names no specific crawler and offers no measurement of how often the rules are skipped.

Coding agents reading other people's repositories

A section headed "Coding-Agenten holen Trojaner" — coding agents fetch trojans — describes the second pattern: AI agents crawl Git projects without being invited. The objection cuts both ways. Pulling in unvetted third-party code and running it is also a route for malware to reach the machine doing the pulling.

Books bought in order to be destroyed

Third comes the physical supply chain: AI firms buying old books by the ton, shipping them across oceans, scanning them and throwing the remains away. No company and no volume figure accompany the claim. The column separately reports that the publisher Carlsen is suing OpenAI over imitation of the style of Astrid Henn's Neinhörner books.

Inside the consulting room

The sharpest example sits under the heading "KI lauscht beim Arzt," AI eavesdrops at the doctor's. Citing reporting by Zeit, Hosbach writes that Doctolib wants to use the content of recorded doctor-patient conversations for AI training. Consent is collected in advance, he notes, but through the app and in an unrelated context. Medical conversation is among the most strongly protected categories of speech there is.

What this piece does not establish

Only the heise column itself was available for this report. Its individual cases — Doctolib, the Carlsen suit, the bulk book purchases — carry no case numbers, sums or named vendors in the original, and none of them were independently verified here. Read them as claims made inside a commentary.

Hosbach's own conclusion is defensive rather than regulatory: look after your own data, because for the moment anything attached to a cable belongs to the crawlers.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

Who wrote this commentary and where did it appear?

Wolf Hosbach, published September 5, 2026 in the opinion section of heise online as an installment of the WTF column.

Does robots.txt actually stop AI crawlers?

Not technically. It is a voluntary convention, so a crawler that ignores it still reaches the content — which is precisely the column's accusation.

Is Doctolib training AI on recorded doctor visits?

The column cites reporting by Zeit that the company intends to, gathering consent through its app in a different context. We have no independent confirmation of that claim.