Harvey's Margins Go From -50% to Positive on Kimi K3
A.I. / news
Harvey's Margins Go From -50% to Positive on Kimi K3
The $15.5 billion legal AI startup's token costs rose twentyfold under OpenAI and Anthropic's usage pricing, and Bloomberg reports Abridge, Decagon and Ramp are making the same open-weight switch.
The short version
Harvey's gross margin fell from about 50 percent at the start of 2026 to roughly negative 50 percent by June, then turned positive again in August after the legal AI startup shipped its own model instead of buying inference from OpenAI and Anthropic, Bloomberg reported Sunday. The new model, Tenet, is built on Moonshot AI's open-weight Kimi K3 rather than a frontier lab's API. Harvey, valued at $15.5 billion and backed by the OpenAI Startup Fund, Sequoia Capital and Andreessen Horowitz, did not change its prices to customers.
What broke the margin in the first place
A March agent update drove a spike in customer usage, and Harvey's token consumption rose twentyfold over the year under OpenAI and Anthropic's usage-based enterprise pricing, according to Bloomberg and a separate account from newsletter The Daily Brief. Running more agent turns per matter meant paying frontier-lab rates on volume that had not existed when Harvey priced its contracts, and gross margin went negative by June even as revenue grew.
| Point in time | Harvey's gross margin |
|---|---|
| Start of 2026 | about 50 percent |
| June 2026 | about -50 percent |
| After Tenet launched, August 2026 | positive |
Tenet, the model that fixed it
Harvey and infrastructure partner Fireworks AI post-trained Tenet on a Kimi K3 base using asynchronous reinforcement learning, Harvey said in an Aug. 20 blog post announcing the model as a research preview. Trained on roughly 150 Nvidia B300 GPUs over two months across 1,750 agentic legal task environments, Tenet completes nearly twice as many held-out tasks on Harvey's LAB benchmark as the base Kimi K3 model, and Harvey said it cuts the cost of some review workflows by roughly 90 percent per query.
Harvey is not alone
Bloomberg named Abridge, Decagon and Ramp as pursuing the same shift away from frontier-lab APIs, with Rogo and Canva also weighing versions of the move, according to The Daily Brief. Abridge, a healthcare transcription company, is building a clinical-work model on an open-source Nvidia base; Sequoia Capital and General Catalyst back several of the startups making the switch. None of the other companies has disclosed margin figures as specific as Harvey's.
The pivot puts Harvey in an odd position relative to Anthropic's own findings about Moonshot: the same company whose Kimi model Harvey now runs in production was accused by Anthropic last week of routing more than 23 million customer exchanges to Claude through accounts made to look Singaporean and Japanese. Harvey's blog post does not mention the dispute, and a Kimi K3 license does not depend on Moonshot's own conduct toward its competitors.
What is still unverified
Harvey's improvement figures, the near-doubling of held-out task completions and the 90 percent cost cut, come from its own blog post rather than an outside audit, and the company has not published Tenet's win rate against the frontier models it replaced on the same docket types. Bloomberg's reporting on the margin swing is attributed to the company rather than to Harvey's financial statements, which are private. The next public signal will be whether Harvey's Series F backers disclose updated unit economics, and whether DeepSeek's own price cuts on its comparable V4.1-Flash model change the calculation again before Harvey's next contract renewal cycle.
Sources
More in A.I.
- 01Xiaomi's MiMo-V2.6-Pro Matches Grok 4.7 for $2.62 MillionThe MIT-licensed, trillion-parameter model tops Artificial Analysis' open-weight ranking and beats DeepSeek's V4.1-Flash on the same index, though Xiaomi's own numbers show it still trails Claude Opus 5 on some tasks.
- 02GPT-6 Astra Refuses Just 2 of 100 Unsafe Robot CommandsRobocurve's RoboHarm benchmark had Claude Fable 5.1 refuse ten times as often, but rival MolmoAct2's zero refusals came from failing to act, not from restraint.
- 03Bessent Blames OpenAI, Not Agents, for Hugging Face BreachThe Treasury secretary's Monday CNBC remarks reject the frontier labs' push for a liability shield, days after Hugging Face's own account of the July intrusion described one agent, not the 1,200 OpenAI has disclosed.
- 04PrismML Shrinks a 27-Billion-Parameter Model to 5.9 GigabytesTernary-Bonsai-2-27B rewrites Alibaba's Qwen3.8-27B in three-value weights, keeping 98.2% of its benchmark score at roughly a ninth of the size, PrismML said.