Harvey Tenet runs on Kimi K3. The benchmark came first
Ricardo Argüello, September 15, 2026
CEO & Founder
General summary
On August 20, 2026 Harvey released Tenet, its first in-house post-trained model, built on Moonshot AI's open-weight Kimi K3 together with Fireworks. It takes the top score on LAB Contracts and second place on LAB, the legal agent benchmark Harvey itself publishes.
- Tenet completes almost twice as many held-out LAB tasks as base Kimi K3, and 20% more on LAB Contracts
- All-pass rate rises 9 percentage points on LAB and 2 on LAB Contracts
- First on LAB Contracts, second on LAB, which is a narrower claim than the coverage made
- Gains transfer to Mercor's APEX Agents and Crosby's Redline Bench, neither seen during training
- Harvey previously routed customer work through models from OpenAI, Anthropic and Google
Think of a racing team that leased its engine for years. Every lap billed, every redesign decided by someone else. Harvey bought a second-hand engine block, machined the head to fit its own circuit, and now wins the corner that pays. What made that possible was owning a circuit measured to the centimeter.
AI-generated summary
Harvey has been publishing legal benchmarks for two years. BigLaw Bench, then LAB, the Legal Agent Benchmark. Slow, unglamorous work, mostly read by other people building legal AI.
That is the reason it could ship a model last month.
On August 20 Harvey released Tenet, post-trained on Moonshot AI’s open-weight Kimi K3. First place on LAB Contracts. Second on LAB overall.
Second, not first. The company said so itself, and most of the coverage rounded it up into a story about a Chinese model beating the American labs. The honest number is the interesting one, because it tells you what kind of claim this is.
What the report actually says
Harvey’s write-up reports that Tenet completes almost twice as many held-out tasks on LAB as the base Kimi K3, and 20% more on LAB Contracts. All-pass rate climbs 9 percentage points on LAB and 2 on LAB Contracts.
Then the part that carries the weight. Those gains transfer to two benchmarks Harvey does not own and the model never saw in training, Mercor’s APEX Agents for corporate law and Crosby’s Redline Bench.
A model that only improves on its own benchmark has learned the exam. A model that improves on somebody else’s has learned the work. That distinction is the whole difference between a press release and a result.
The post-training ran with Fireworks, and the method is expensive in the way that matters. Harvey put lawyers on inventing mock disputes and case files, then grading how well models reasoned through them. Human experts, hand-producing evaluation data.
Owning the evaluation is the precondition
Any software company can download Kimi K3 this afternoon. The weights are public and the tooling is commodity.
Almost none of them can answer the next question, which is whether the tuning worked. Answering it takes a body of tasks from your own industry, graded by people who know the industry, large enough and honest enough to separate a real gain from a model that learned to sound more confident.
Without that, post-training is spending money into a hole.
Harvey had the instrument before it had the ambition. Measure first, train second. Reversing that order is how most enterprise fine-tuning projects end up with a model everyone agrees feels better and nobody can defend in a review.
We came at the same idea from the other side in what cannot be trained, which asks which part of your judgment never made it into the data. This is the follow-on. If you can write that judgment down as an exam, it becomes an asset no vendor gets to price.
The Chinese base matters less than the license
There is a geopolitics conversation around this and I understand why it exists, but for an architecture decision it is mostly noise. The weights are open. What serves customers is Harvey’s tuned copy on Harvey’s infrastructure.
The license is the part worth reading. Moonshot requires a separate agreement for model-as-a-service operators past $20 million of revenue in twelve months. Harvey is well past $350 million annualized, so that conversation happened.
Check the license before you fall in love with the model. It is boring and it is where half of these projects quietly die.
The underlying pattern showed up in voice first, when Fish Audio undercut the market by 70% and ElevenLabs grew anyway. Cheapening the model relocates the value upward, into the product and into the evaluation data.
If you sell vertical software
Three questions I would put to your team this week, and I mean them literally.
How many real customer tasks do you have graded by someone qualified to grade them? Is your pass rate judged by an expert or by another model? What did you spend last month on API calls for work that repeats almost identically?
If the first answer is zero, an open model is not yet useful to you. The weights are fine. You just have no way to find out whether they helped.
That is where our Technology Partner work with software companies starts, building the evaluation set for your vertical before touching a single model weight. Harvey took two years to get theirs. You do not need two years. You need to start at the same end.
Build the benchmark for your verticalFrequently Asked Questions
Harvey Tenet is the first in-house post-trained model from legal AI company Harvey. It is built on Kimi K3, the open-weight model released by Chinese startup Moonshot AI in July 2026, and was post-trained together with Fireworks for long-horizon legal agent work.
Harvey reports that Tenet completes almost twice as many held-out LAB tasks and 20% more LAB Contracts tasks than the base Kimi K3 model, with all-pass rate rising 9 and 2 percentage points. It places first on LAB Contracts and second on LAB overall.
LAB is the legal agent benchmark Harvey builds and publishes alongside BigLaw Bench. It measures model performance on real long-horizon legal tasks rather than legal reasoning in structured formats. Harvey uses it both to evaluate outside models and to train and validate its own.
To turn a variable cost into a fixed one and to control the improvement cycle. Harvey previously routed customer work through OpenAI, Anthropic and Google models on a per-call basis. Owning post-trained weights means the company decides when it tunes and on which data.
Related Articles
Google Entered Legal and Harvey Became a Connector
Gemini Enterprise for Legal shipped August 25 with Cleary, Freshfields and Weil. Harvey appears inside it as an MCP integration rather than a rival.
FairMindSim: An Inverted U Built From Just 10 Models
A KDD 2026 paper says mid-tier models punish twice as hard as humans. The curve rests on 10 models and a correlation that misses conventional significance.
No US Ban on Chinese AI Models Yet, But the Mechanism Is Loaded
Washington has not banned Chinese AI models. But the executive order and Entity List addition are already drafted, ready without a Congressional vote.
Jensen Huang Said It in Two Words: Your Moat Is Knowing More
Jensen Huang confirms that AI competitive advantage isn't about models, it's vertical specialization. What that means for B2B company leaders.
Qwen3.8-Max redesigned a chip from 8,298 gates to 678
Alibaba opened the weights of a 2.4T-parameter model that ran ~500 closed-loop chip design iterations alone. Your evaluation criteria just went stale.
Jensen Huang Authored a Letter. Anthropic Didn't Sign
77 companies signed Jensen Huang's letter backing open-weight AI. Anthropic and Amazon didn't. The real reason matters to marketing teams, not just IT.