Skip to main content

Thomson Reuters Trained Its Own Model for $450,000

Thomson Reuters built a legal model that beats GPT 5.4 and Claude Sonnet 5 on its own content. The final training run cost $450,000. The corpus is the moat.

Thomson Reuters Trained Its Own Model for $450,000

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

AI & Automation 4 min read

The number that matters in this story is $450,000.

That is what the final training run cost for Thomson, the legal language model Thomson Reuters launched on August 24.

Less than a mid-sized enterprise spends annually on seats nobody logs into.

The corpus won, not the architecture

Thomson was not built from scratch. It started from an open-weights model, most recently Qwen 3.5, and got trained on material nobody else can license: Westlaw, Practical Law, Checkpoint, the Reuters archive.

Under 10% of the company’s total proprietary content. Zero customer data.

By Thomson Reuters’ own evaluations, with access to that content the model outperformed GPT 5.4 and Claude Sonnet 5 on legal research. On tests using only web information it landed roughly even, and the company says plainly it is not the leader there yet.

The gap between those two conditions is the entire finding. They did not win on technique or scale. They won at home, cooking with ingredients only they can buy.

CTO Joel Hron framed it as evidence that there is another path than treating scale as the answer, which is a careful way of saying the frontier labs do not automatically own every vertical.

Read the $40 million correctly

The figure that will get quoted is $40 million over two years in talent and compute.

The figure that teaches you something is the ratio. If the final run was $450,000, then the $40 million went to people, failed experiments, data curation, and the slow work of deciding what a correct answer looks like in law. That last one is not an engineering question and no amount of compute settles it.

The dominant cost was encoding judgment.

Which reframes the whole build-or-buy conversation. The gating question is not whether you can afford GPUs. It is whether you own a corpus and employ people who can tell a good answer from one that merely sounds good.

We watched the same shape when Bridgewater tuned a model and beat the frontier labs. The pattern repeats with a regularity that is starting to get boring.

You probably should not train a model

Here is the unpopular half. Almost no mid-sized company should be training its own model in 2026.

Thomson Reuters holds a legal archive assembled over more than a century, sells exactly that for a living, and employs people who can grade a legal answer. Most companies have partial versions of all three.

The lesson still transfers, at a different scale. Your edge is the body of knowledge only you hold: how you price, why you lost the last twenty deals, which exception applies to which account, what actually broke in the March rollout.

None of that requires training anything. It requires the knowledge to exist somewhere queryable and structured, instead of in email threads and three people’s heads. We made that case in the model is your database.

Where I would start on a Thursday

Pick one task your team does often that requires real judgment. Quoting. Reviewing a contract clause. Deciding whether an unusual order gets approved.

Collect twenty resolved cases with the correct answer and the reason it was correct. Twenty, not two hundred.

Then measure a commercial model on that task twice: once with those twenty cases as context, once without. The delta between those two runs is your advantage, expressed as a number, and you can have it by dinner.

Thomson Reuters spent $40 million learning this at industrial scale. Your version costs an afternoon and two people, and answers the same question: is the value in the model, or in what you know?

Let’s find where your data advantage actually is

Frequently Asked Questions

Thomson Reuters specialized models proprietary data Westlaw Qwen legal AI B2B strategy

Related Articles

Qwen3.8-Max redesigned a chip from 8,298 gates to 678
Business Strategy
· 8 min read

Qwen3.8-Max redesigned a chip from 8,298 gates to 678

Alibaba opened the weights of a 2.4T-parameter model that ran ~500 closed-loop chip design iterations alone. Your evaluation criteria just went stale.

Qwen Alibaba open weights
Blackstone bet on Norm AI. The market just agreed.
Business Strategy
· 7 min read

Blackstone bet on Norm AI. The market just agreed.

Blackstone put $50M into Norm AI in November and started using it inside its own legal function. Eight months later, a Series C valued the company at $1.2B.

Norm AI Blackstone legal AI
Porsche Sold Its Consulting Arm, Then Rented It Back
Business Strategy
· 4 min read

Porsche Sold Its Consulting Arm, Then Rented It Back

Porsche sold MHP to TCS for 320 million euros and signed a five-year partnership worth 1.25 billion. The accounting works. The learning walks out.

Porsche TCS MHP
OpenAI Cuts Cursor Off Over a Change-of-Control Clause
Business Strategy
· 4 min read

OpenAI Cuts Cursor Off Over a Change-of-Control Clause

OpenAI is ending Cursor's model access on November 12 because SpaceX bought it. Cursor survives because OpenAI was only 5% of its traffic. What is your number?

OpenAI Cursor SpaceX
Perplexity at $30B on $750M, and Nvidia May Fund It
Business Strategy
· 4 min read

Perplexity at $30B on $750M, and Nvidia May Fund It

The Information reports Nvidia in talks to back Perplexity above $30B. In January it committed its entire annual revenue to three years of Azure compute.

Perplexity Nvidia answer engines