Thomson Reuters builds its own AI model for $40 million
The information giant launches "Thomson", its own Qwen-based language model — instead of continuing to rent from OpenAI or Anthropic.

Illustration · AI-generated (AI IN LIFE)
At a glance
- Unveiled: 24 August 2026 (press release, The Decoder, SiliconANGLE)
- Investment: about $40M over two years; final training run about $450,000
- Base: Alibaba's Qwen3.5-397B, retrained with Imperial College (intermediate version "Snowdon")
- Benchmark: behind GPT-5.5 and Gemini 3.1 Pro on LegalBench — ahead only with Westlaw data access
- Goal: independence from providers like OpenAI and Anthropic, protection of proprietary data
Thomson Reuters has unveiled "Thomson", its own language model — framing the move as independence from the big AI providers. The company says it invested around $40 million over two years in staff and compute; the final training run itself cost only about $450,000, as The Decoder reports.
Technically, Thomson builds on Alibaba's open Qwen3.5-397B model and was retrained in partnership with Imperial College — including an intermediate version called Snowdon for safety and ethics fine-tuning. The real lever is data: the model taps exclusive content such as Westlaw and Practical Law.
CTO Joel Hron describes the strategy as recognizing "which intelligence is important enough to own." Evaluation lead Andrew Bean stays sober: without the proprietary data, Thomson is "within the scope of the other models, but certainly not the leader yet."
Benchmarks confirm it: on Stanford's LegalBench, Thomson trails GPT-5.5 and Gemini 3.1 Pro. It leads only where it can access the exclusive legal databases — which, from the company's perspective, is exactly the point: value comes from data no competitor can license.
The signal extends beyond legal. If a data company can build a usable specialist model on an open-source base for $40 million — a fraction of frontier budgets — then "own instead of rent" becomes a real option for many businesses with valuable data assets. Pressure on the API business models of OpenAI and Anthropic is growing.
FAQ
Why is Thomson Reuters building its own model?
To avoid vendor lock-in, protect its data and capture the value of exclusive content like Westlaw itself — instead of renting models indefinitely.
Is Thomson better than GPT-5.5?
No. It trails on general benchmarks like LegalBench; it only wins in combination with the company's exclusive legal databases.
What does this mean for other companies?
Specialist models on open-source bases are becoming affordable. Companies with valuable proprietary data can increasingly use it in their own model rather than through third-party APIs.


