Popular AI models give incorrect financial advice on average in 57% of cases, and for complex questions the share of errors reaches 88%. These are the conclusions reached by fintech company Saturn after checking the responses of 18 popular models, including ChatGPT, Claude, Copilot, Grok, and Gemini.

The study shows that the problem is not limited to occasional factual mistakes. Neural networks can miscalculate amounts, overlook important risks, fail to account for changes in tax rules, and even invent non-existent regulations. At the same time, such answers may appear convincing enough for a user to mistake them for professional financial advice.

More than 10,000 answers to financial questions

Saturn published a study titled “Artificial Authority: Can AI Be Trusted to Provide Financial Advice?”. The authors tested 18 models by asking them 121 financial questions. Each question was asked five times, and as a result the researchers collected more than 10,000 responses.

On average, 57% of the advice turned out to be incorrect. At the same time, the results varied significantly depending on the type of model. Free versions were wrong in 63% of cases, while for paid ones this figure was 49%.

On complex financial questions, the situation became even worse. Free models gave incorrect answers in 93% of cases, and for some systems the error rate reached 99%.

The researchers identified several recurring types of mistakes: incorrect calculations, lack of warnings about financial risks, ignoring upcoming changes in tax legislation, and references to rules that do not actually exist.

An AI mistake could cost a user thousands of dollars

Some examples cited in the study show how serious the consequences of such answers can be.

In one case, Claude Haiku 4.5 gave incorrect advice on retirement tax. According to the study’s authors, if a user had followed that advice, they could have faced a fine of $23.3 thousand.

In another case, AI advised a person with debts to stop paying rent and council tax in order to direct the money toward repaying a loan with a high interest rate. Such advice could have led to eviction and bailiff involvement.

The mistakes also concerned student loans. Claude stated that a graduate could move to another country and avoid repaying the debt. According to the study, in practice this would not have freed the borrower from their obligations and could have led to an increase in their monthly payments.

Another example involved a mortgage. Gemini incorrectly told a borrower that a payment holiday would not affect their credit score. The study’s authors note that such information could lead to less favorable lending terms in the future.

People are already using AI for financial decisions

The problem Saturn points to is not only the accuracy of individual answers. The more often people turn to general-purpose chatbots for financial recommendations, the greater the potential consequences of such mistakes.

Saturn CEO Amal Jolly said that the poor quality of advice from mass-market AI models could lead to large-scale harm for consumers: people trust systems that are capable of giving convincing answers even when those answers are wrong.

As the British publication The Intermediary notes, Saturn’s study was released shortly after the publication of the Mills Review by the UK Financial Conduct Authority (FCA). According to that review, 26% of consumers already turn to general-purpose AI models for money-related questions.

Thus, the problem is not only that chatbots sometimes make mistakes. A user may fail to recognize the error and make a decision about taxes, credit, a mortgage, or debt repayment based on an answer that sounds confident but does not take into account specific circumstances or current rules.

Saturn’s study does not mean that AI is useless for working with financial information. However, the results show a gap between models’ ability to quickly explain financial issues and the reliability of their recommendations when a real financial decision depends on the answer.