Skip to main content

Harvey Tenet runs on Kimi K3. The benchmark came first

Harvey stopped renting frontier models and post-trained an open Chinese base. It placed first on contracts, second overall, on a benchmark it owns.

Harvey Tenet runs on Kimi K3. The benchmark came first

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

Business Strategy 4 min read

Harvey has been publishing legal benchmarks for two years. BigLaw Bench, then LAB, the Legal Agent Benchmark. Slow, unglamorous work, mostly read by other people building legal AI.

That is the reason it could ship a model last month.

On August 20 Harvey released Tenet, post-trained on Moonshot AI’s open-weight Kimi K3. First place on LAB Contracts. Second on LAB overall.

Second, not first. The company said so itself, and most of the coverage rounded it up into a story about a Chinese model beating the American labs. The honest number is the interesting one, because it tells you what kind of claim this is.

What the report actually says

Harvey’s write-up reports that Tenet completes almost twice as many held-out tasks on LAB as the base Kimi K3, and 20% more on LAB Contracts. All-pass rate climbs 9 percentage points on LAB and 2 on LAB Contracts.

Then the part that carries the weight. Those gains transfer to two benchmarks Harvey does not own and the model never saw in training, Mercor’s APEX Agents for corporate law and Crosby’s Redline Bench.

A model that only improves on its own benchmark has learned the exam. A model that improves on somebody else’s has learned the work. That distinction is the whole difference between a press release and a result.

The post-training ran with Fireworks, and the method is expensive in the way that matters. Harvey put lawyers on inventing mock disputes and case files, then grading how well models reasoned through them. Human experts, hand-producing evaluation data.

Owning the evaluation is the precondition

Any software company can download Kimi K3 this afternoon. The weights are public and the tooling is commodity.

Almost none of them can answer the next question, which is whether the tuning worked. Answering it takes a body of tasks from your own industry, graded by people who know the industry, large enough and honest enough to separate a real gain from a model that learned to sound more confident.

Without that, post-training is spending money into a hole.

Harvey had the instrument before it had the ambition. Measure first, train second. Reversing that order is how most enterprise fine-tuning projects end up with a model everyone agrees feels better and nobody can defend in a review.

We came at the same idea from the other side in what cannot be trained, which asks which part of your judgment never made it into the data. This is the follow-on. If you can write that judgment down as an exam, it becomes an asset no vendor gets to price.

The Chinese base matters less than the license

There is a geopolitics conversation around this and I understand why it exists, but for an architecture decision it is mostly noise. The weights are open. What serves customers is Harvey’s tuned copy on Harvey’s infrastructure.

The license is the part worth reading. Moonshot requires a separate agreement for model-as-a-service operators past $20 million of revenue in twelve months. Harvey is well past $350 million annualized, so that conversation happened.

Check the license before you fall in love with the model. It is boring and it is where half of these projects quietly die.

The underlying pattern showed up in voice first, when Fish Audio undercut the market by 70% and ElevenLabs grew anyway. Cheapening the model relocates the value upward, into the product and into the evaluation data.

If you sell vertical software

Three questions I would put to your team this week, and I mean them literally.

How many real customer tasks do you have graded by someone qualified to grade them? Is your pass rate judged by an expert or by another model? What did you spend last month on API calls for work that repeats almost identically?

If the first answer is zero, an open model is not yet useful to you. The weights are fine. You just have no way to find out whether they helped.

That is where our Technology Partner work with software companies starts, building the evaluation set for your vertical before touching a single model weight. Harvey took two years to get theirs. You do not need two years. You need to start at the same end.

Build the benchmark for your vertical

Frequently Asked Questions

Harvey Kimi K3 Moonshot AI open weights vertical specialization benchmarks legal software

Related Articles

Google Entered Legal and Harvey Became a Connector
Business Strategy
· 4 min read

Google Entered Legal and Harvey Became a Connector

Gemini Enterprise for Legal shipped August 25 with Cleary, Freshfields and Weil. Harvey appears inside it as an MCP integration rather than a rival.

Google Cloud Gemini Enterprise Harvey
FairMindSim: An Inverted U Built From Just 10 Models
Business Strategy
· 7 min read

FairMindSim: An Inverted U Built From Just 10 Models

A KDD 2026 paper says mid-tier models punish twice as hard as humans. The curve rests on 10 models and a correlation that misses conventional significance.

FairMindSim KDD 2026 AI moral judgment
Jensen Huang Said It in Two Words: Your Moat Is Knowing More
Business Strategy
· 6 min read

Jensen Huang Said It in Two Words: Your Moat Is Knowing More

Jensen Huang confirms that AI competitive advantage isn't about models, it's vertical specialization. What that means for B2B company leaders.

AI strategy vertical specialization competitive advantage
Qwen3.8-Max redesigned a chip from 8,298 gates to 678
Business Strategy
· 8 min read

Qwen3.8-Max redesigned a chip from 8,298 gates to 678

Alibaba opened the weights of a 2.4T-parameter model that ran ~500 closed-loop chip design iterations alone. Your evaluation criteria just went stale.

Qwen Alibaba open weights
Jensen Huang Authored a Letter. Anthropic Didn't Sign
AI in Marketing
· 5 min read

Jensen Huang Authored a Letter. Anthropic Didn't Sign

77 companies signed Jensen Huang's letter backing open-weight AI. Anthropic and Amazon didn't. The real reason matters to marketing teams, not just IT.

open weights Anthropic Jensen Huang