What is Harvey Tenet in 2026: Kimi K3 post-training for legal agents
Harvey named its first post-trained open-weight model Tenet. This piece separates the official post from Law.com’s 20 August recap.
This week the legal-tech firm Harvey put a research preview into a checkable technical note: instead of only renting someone else’s flagship API, it post-trained an open-weight base for legal work and published numbers on its own Legal Agent Bench.
The object is Harvey Tenet. The official post calls it a Kimi K3 base, asynchronously reinforcement-learned with Fireworks, aimed at long-horizon agentic legal tasks. What follows is that post plus Law.com on 20 August — not a claim that Tenet now beats Fable 5 or GPT-5.6 across the board.
What the official note says
Harvey writes two six-month aims: frontier legal intelligence on open weights, and systems so firms can build and own models. Tenet is the first post-trained checkpoint on the first line. Training environments copy LAB: a partner-style request, a client matter, an expert rubric. The agent works in a sandbox, writes deliverables to disk, and is graded by an LLM judge.
Figures that can be checked
- 1
The base and the optimizer are named
Official: Kimi K3 + Fireworks async RL; policy via GSPO (in-group advantages, rejudge near-ties, a length term, double-sided clipping). Judge ablations settled on Kimi 2.6. Data: synthetic, public legal, and human-expert sets, with Mercor and others on expert data and synthetic review.
- 2
LAB numbers are versus K3, not an outside lab board
Official: almost 2× hold-out completions, +20% on LAB Contracts; all-pass +9 and +2 pp. It claims first on Contracts, second on LAB. A figure notes Vals LAB as the baseline source. No third-party rerun of the full table was published.
- 3
Three other capability runs are not the Tenet checkpoint
M&A diligence describes GLM-5.2 in an RLM harness after self-distillation, at 60.1% criteria pass. Review Table is also a post-trained GLM-5.2. Firm Knowledge is a Qwen3.8-27B with Engram. Those are add-on tracks “to be incorporated,” not Tenet itself.
Cost is a Pareto claim, not a public price list
The post says open weights are cheaper per token and that reward shaping cut inference tokens. It shows a quality–cost plot on LAB hold-out. It does not publish a per-million-token list price in the essay.
“Open-weight” is not the same as a public download
The official title and Law.com both say open-weight. The same page marks a Research Preview and says research still has to reach production. A repo URL and full license were not written as a public weights drop.
Valuation figures conflict; treat them as press
Startup Fortune wrote a valuation near $15.5 billion; Business Insider wrote an $11 billion legal-software business. Both are media lines. What can be checked is the Tenet methods note, not a new funding round.
| Claim | Checkable source | How to read it |
|---|---|---|
| K3 base + Fireworks async RL | Harvey official post | Methods claim on a research preview |
| First on LAB Contracts, second on LAB | Harvey post; Law.com recap | Vendor eval; comparators in Law.com |
| No customer data | Harvey training section | Company statement, not an audit |
The post says gains transfer to APEX Agents and Redline Bench, unseen in training. That is their generalization claim. The benches may still overlap legal work; it is not a court certification.
# Harvey Tenet (official, research preview)
base: Moonshot Kimi K3
post-train: async RL + Fireworks
optimizer: GSPO
judge: Kimi 2.6
no customer data in post-trainingHow to read the boundaries
- 2× completions ≠ unsupervised legal opinions
- The figures are LAB hold-out plus rubrics. The post also says knowledge benches (LegalBench, CUAD, MAUD) did not hurt the base. That is not a license to practise or a regulator’s approval.
- An LLM judge is noisy
- Reward mixes fine rubric terms, legal issues solved, and a perfect-score bonus. They picked Kimi 2.6 after ablating heavier flagships. Judges err.
- Harvey II Memory is a product, not Tenet
- Law.com: Harvey II and retained context landed about two days before Tenet. Startup Fortune wrote Memory is not used to train models. Read them apart.
Questions worth checking
Has Tenet already replaced OpenAI / Anthropic for every client?
The official post does not say that. It is a research preview; production compute is listed as next. Harvey’s business was built on third-party models, as Law.com and Business Insider both recount.
Is “first on LAB” an independent-lab finding?
No. LAB is Harvey’s own open-source legal-agent bench. The company claims first on Contracts and second overall. Law.com named comparators. A third-party rerun was not shown.
Is the 60.1% diligence figure Tenet’s score?
No. That section describes GLM-5.2 + RLM after self-distillation, with a technical report still to come. Do not stack it on Tenet’s LAB numbers.