Product · September 22, 2026
AI chatbots fail on over half of financial questions
A recent survey by Saturn Technology examined how well large language models answer financial queries. The study tested twelve major models, including ChatGPT, Claude, Copilot, Grok and Gemini, by posing 121 questions five times each, yielding more than 10,000 responses. Errors occurred on 57 percent of all questions, rising to 88 percent for the most complex queries that required multiple calculations. Free versions made mistakes on 63 percent of simple questions, while paid tiers erred on 49 percent. The top performer, Claude Opus 5 Reasoning, still failed on 39 percent of easy items and on 67 percent of difficult ones.
Saturn concluded that every model showed low accuracy across all categories. Investment director Robert Næss of Nordea commented that the findings highlight a need for clearer questioning and greater context when using AI for finance. He noted that a test question about prioritising debt repayment received a technically correct but overly simplistic answer. Another question concerning student loans in the UK produced an incorrect claim that Plan 1 loans are automatically cancelled when a borrower moves abroad; Saturn clarified that repayment remains obligatory.
The report warns that following AI advice can lead to severe financial loss, legal trouble or even life‑threatening outcomes. Næss suggested that Saturn’s commercial interests may have influenced how the test was framed. He urged users to verify AI outputs and to provide precise context, such as specifying weather conditions for ski plans. Subscriptions for higher‑performing models start at 64 kroner per month, with premium versions costing around 200 kroner monthly.
Reported by E24.