Saturn Finds AI Models Wrong on Finance 57% of the Time
A.I. / news
Saturn Finds AI Models Wrong on Finance 57% of the Time
Claude Opus 5 was the most accurate of 18 chatbots Saturn tested and Claude Haiku 4.5 the least, with one pension error risking a £17,500 HMRC bill.

British fintech firm Saturn tested 18 AI models against 121 financial questions and found wrong answers 57% of the time, according to a report the company published on Sept. 14. Each question was asked five times to check consistency, producing more than 10,000 graded responses in total.
Saturn scored an answer as a failure if it contained a factual error, missed a key point, or left out an important warning, Financial Reporter UK reported. The questions covered debt management, mortgages, pensions, tax and student loans. On the hardest questions, the average error rate rose to 88%, and some models were wrong 99% of the time.
Claude Opus 5 scored best, Claude Haiku 4.5 scored worst
Saturn's per-model breakdown ranged from a 39% error rate to 82%, spanning five major AI labs' products in the same test.
| Model | Error rate |
|---|---|
| Claude Haiku 4.5 | 82% |
| Google Gemini 3.1 Pro | 73% |
| xAI Grok 4.5 | 59% |
| ChatGPT 5.6 Luna | 58% |
| Claude Opus 5 (reasoning) | 39% |
Free tools were wrong 63% of the time, against 49% for paid tools, and free models failed 93% of the hardest questions, Financial Reporter UK reported.
- Free models63 %
- Paid models49 %
Source: Saturn, "Artificial Authority" report, accessed 2026-09-21
A pension error, priced at £17,500

One mistake on pension tax rules could have left a consumer facing a £17,500 charge from HM Revenue and Customs, according to Saturn's findings. Other errors included a model recommending a consumer pay off their highest-interest debt before priority bills, which Financial Reporter UK said risked eviction or legal action, and a model that falsely told a user student loan repayments could be halted by moving abroad.
"The low quality of financial advice from mainstream AI models risks leading to widespread consumer harm," Saturn Chief Executive Amal Jolly said in comments reported by Financial Reporter UK.
Saturn sells compliance software to the advisers it is comparing AI against
Saturn is not a neutral tester. The company builds AI and workflow software for human financial advisers, serving more than 500 advice firms, and a finding that general-purpose chatbots give unreliable financial advice supports the case for the compliance-focused tools Saturn sells. Professional Adviser also covered the 57% figure on Sept. 14, and neither outlet reported an attempt by Saturn to have the methodology checked by an outside lab.
That commercial interest does not make the underlying numbers wrong, but Saturn's benchmark is Saturn's own; no independent party has re-run its 121 questions against the same 18 models. The pattern recalls other vendor-run comparisons this site has covered, including DeepSeek's V4.1-Flash, which topped its own charts before an outside lab found a gap between the vendor's claims and independent results. Saturn's report itself argues the fix is guardrails on the models, not blanket regulation of AI in finance, a framing that also favors a compliance-software vendor over an outright ban. The same pattern showed up when a startup's model launch ran into a researcher's prior-art claim: a vendor's own numbers set the terms of the story until someone outside the company checked them.
Saturn did not say whether it plans to re-test the same 18 models as new versions ship, or whether it will publish the full 121-question set so another lab can attempt to reproduce the 57% figure.
Sources
More in A.I.
- 01Qwen-Image-2.1 Ships With Native TransparencyThe 7-billion-parameter model generates and edits RGBA images in one pass, but Alibaba's research licence bars commercial use without a separate grant.
- 02OpenAI Sets Three-Track System for Misalignment ReportsThe company published six reports under the new process on Sept. 17, including one where a model searched GitHub for a leaked key and then invented the data it could not retrieve.
- 03Cloudflare Security Tool Gains 2,428 Stars in a DayThe three-month-old skill topped GitHub's trending list Monday, and Cloudflare's own numbers show a single run still misses about half the bugs it is looking for.
- 04China Telecom's Xing4.0 Trains Entirely on Huawei ChipsThe 29-billion-parameter Xing4.0-29B-A4B activates 4 billion parameters per token and scores 75.00 on SWE-bench Verified, a point behind Qwen3.6-35B-A3B's 76.00.