AI financial advice is solid but quietly favours confident users
Researchers at MIT and Stanford simulated entire lifetimes of following chatbot money advice. It mostly matched what finance textbooks recommend — but the same systems steered women and less experienced users towards tens of thousands of dollars less by retirement.

A team of finance researchers at MIT Sloan and Stanford decided to test something a large slice of the public is already doing without much scrutiny: typing a question about money into a chatbot and then acting on whatever comes back. Their verdict, set out in a paper that took the Swiss Finance Institute's Outstanding Paper Award for 2026, is that the machines dispense guidance that is broadly sensible — sensible enough that almost anyone over 30 who followed it would end up with a real cushion of savings.
The complication is that the same systems produce measurably different advice depending on who appears to be asking, and those differences compound over a working life into gaps of tens of thousands of dollars.
Sixty-seven simulated years of doing what the chatbot says
Judging financial advice is hard because the payoff arrives decades later. So the researchers — Taha Choukhmane, Weidong Lin and Matthew Akuzawa at MIT Sloan, with Tim de Silva at Stanford's Graduate School of Business — began by building an economic model of a typical life: how earnings rise and fall, how jobs are lost and found, how investments grow, how tax bites. That model produces a yardstick for what an optimal decision looks like at any given age, against which real advice can be scored.
Then came the human input. A sample of 1,000 adults were asked to write their own questions about spending and investing, in their own words, and put them to three large language models — the systems behind consumer chatbots, trained on enormous quantities of text to predict plausible replies. The models tested were GPT-5.2, GPT-5.6 and Gemini 3 Flash. The researchers then ran the resulting recommendations forward through their simulated lifetime, from age 22 to 89, putting the same sorts of questions to the chatbot again and again as circumstances changed and obeying the answers each time.
Finally they repeated the whole exercise using questions written the way a finance professor would write them: full disclosure of age, employment, income and balances, plus explicit assumptions about life expectancy, retirement age, job risk and the stability of tax and pension rules.
What emerged was better than the authors expected. Across both the amateur and the academic questions, the chatbots pushed people to save while working and to spend those savings down in retirement, to hold shares rather than sit in cash, and to hold them through broad diversified funds rather than individual bets. They also recommended trimming exposure to shares from around age 45 — the standard advice that risk should fall as the years available to recover from a crash shrink. That the models landed on textbook asset allocation, Choukhmane noted, was not a foregone conclusion given the loose and often vague questions people actually typed.
What the models got wrong was everything that required a second thought
The failures clustered in one place: adaptation. The chatbots reached for fixed rules of thumb — save this share of income, hold this fraction in shares — and then applied them rigidly as the simulated life moved on.
The clearest example was job loss. Told that a person had become unemployed, the models advised cutting household spending far harder than the situation warranted, even when that person had a healthy balance sitting in the bank. Savings exist precisely to smooth over such gaps; the advice treated them as untouchable.
The models were also poor at rebalancing — the housekeeping of periodically selling whatever has grown too large a slice of a portfolio and buying back what has shrunk, so the mix of risk stays where it was intended. Left to the chatbots, portfolios simply drifted with the market. Better-structured questions improved the spending and saving guidance considerably, but even the professor-grade prompts did not produce enough active rebalancing.
The same question, worded differently, is worth $50,000
The most uncomfortable finding concerns who benefits. Simulated users whose prompts came from men, from people who scored highly on financial literacy, or from people who had used AI for money questions before ended up roughly 5% wealthier as they approached retirement.
Broken down: the models recommended heavier share allocations to men and to the financially literate, a difference worth about $50,000 — some 4% — by age 60. They recommended lower saving rates to people with no prior experience of asking AI about money, which cost that group nearly $100,000, around 6%, over the same span.
Part of that is simply what people asked. The vocabulary differed sharply: women's prompts were more likely to mention family, groceries and bills, men's to mention strategy, growth and crypto. But roughly a third of the gender gap survived even when the question was held identical and only the stated gender changed — the model altered its answer on the basis of the label alone.
Choukhmane is careful here. Some variation is defensible; women live longer on average and face different income risk, so identical advice would not necessarily be correct advice. The problem is that nobody can currently tell defensible variation from prejudice absorbed from training data, because the field has no agreed benchmark for how good advice ought to differ across demographics. Until it does, the models have nothing to be measured against.
Fund companies are being recommended without asking
One finding should worry marketing departments. The chatbots frequently named specific account types, products and providers that the person had never mentioned. Vanguard funds turned up in 6% of responses and iShares in 3.4%, despite fewer than 0.4% of the questions referencing either firm.
For financial companies, that hints at a shift in how customers are acquired: less about advertising or ranking well in search, more about how a model happens to describe your product when somebody asks it where to put $200 a month.
For everyone else, the practical takeaway is less about obedience than education. The strongest argument for chatbot advice is cost — the people who most need help with retirement planning are usually the ones who cannot pay a human for it. But the study suggests the benefit flows disproportionately to those already equipped to ask a good question. The open problem, as Choukhmane frames it, is making the advice work for people who write imperfect prompts, which is to say nearly everyone.
The full write-up of the study is available from MIT Sloan.
Questions
Does the study say chatbots give better advice than human financial advisers?
Not quite. The researchers found that AI advice avoids the fees, sales incentives and conflicts of interest that come with paid advisers, and that it is genuinely accessible to people who could never afford one. But they frame it as a complement rather than a replacement — useful for putting a professional's plan into practice between meetings, and for building financial understanding rather than for blind obedience.
What makes a prompt about money a 'good' one?
The versions that produced better results spelled out the full picture: age, employment status, income, existing balances, planned retirement age, and explicit assumptions such as normal life expectancy and unchanged tax and pension rules. Vague requests — the study cites the example of asking where to invest $50 plus $25 a month — invite generic rules of thumb.
Which chatbots were tested?
The 1,000 participants put their questions to GPT-5.2, GPT-5.6 and Gemini 3 Flash. The findings describe the behaviour of those models at the time of testing, not of AI assistants in general.
Can users do anything about the gender and literacy gaps?
In the short term, the researchers suggest explicitly instructing the model to guard against bias, and giving it far more context so it has less room to infer things from tone or vocabulary. The longer-term fix is not something users can apply: the field needs an agreed standard for how sound advice should legitimately vary between people.