AI managed household portfolios: A preliminary report
This research investigates AI-managed household portfolios, finding that large language models recommend undiversified portfolios concentrated in high-momentum, large-cap technology stocks. While returns may exceed market benchmarks, they do not provide abnormal returns when adjusted for risk characteristics. Recommendations are primarily driven by media attention rather than investment skill.
Please login or join for free to read more.
OVERVIEW
Introduction
Almost every sector of the economy has experienced marked creative destruction due to the growing use of artificial intelligence (AI). AI now appears to outperform human beings in many settings based on both accuracy and speed. Consequently, there has been rising interest in using AI to manage asset portfolios and extend the capabilities of institutional advisors. Retail investors also have direct access to AI platforms which make financial recommendations without credentials such as FINRA licensing. This warrants an objective academic investigation into AI’s investment style and performance.
This preliminary report addresses several key questions regarding how AI would manage a retail investor’s portfolio. It explores what assets AI recommends, the style of investment, the expected turnover, and how holdings evolve over time. The study tracks whether AI demonstrates stock-picking skill or earns abnormal returns. For nearly a year, the researchers have tracked the performance of AI-managed asset portfolios on a daily basis. The study is entirely prospective, contrasting with other attempts to study AI’s efficacy with backward-looking historical data. The large language models examined include OpenAI’s ChatGPT 5.0 and ChatGPT 5.2, Anthropic’s Claude Sonnet 4.5, Google’s Gemini 2.5 Flash, and xAI’s Grok 4.1 Fast.
Data And Methodology
The data used in this report were collected on a daily basis starting in August 2025. Pre-tests determined that AI responses were invariant to specific language or tone. Prompts fell into two categories: requests for a portfolio of stocks to beat a market index over a one-year horizon (buy-and-hold), and requests to actively manage an existing portfolio (active portfolios). Data collection using buy-and-hold prompts began on 22 August 2025. For active portfolios, ChatGPT 5.0 collection started on 23 September 2025. On 2 January 2026, a prompt was added requesting that the portfolio maintain a market beta between 0.9 and 1.1.
To assess whether AI portfolios generate abnormal returns beyond passive exposure to well-known characteristics, the researchers employed the characteristic-based adjustment of Daniel et al. (1997), known as DGTW. This approach benchmarks each stock against a portfolio of firms with similar size, industry-adjusted book-to-market ratio, and prior-year momentum. This is particularly suited to the setting where AI portfolios exhibit strong tilts toward large-cap, growth, and momentum stocks. Statistical inference was conducted using the estimator by Driscoll et al. (1998) to account for serial correlation and cross-sectional dependence.
Results
The findings indicate that AI chooses high beta stocks when left unconstrained, with an average beta of 1.6. When constrained to a band of 0.9 to 1.1, it persistently recommends portfolios at the higher end. AI portfolios load positively on momentum and negatively on book-to-market and small size. The models recommend holding a small number of assets, exposing users to undiversified idiosyncratic risk. For instance, ChatGPT 5.0 recommends roughly twenty assets, investing 18-20% of portfolio wealth in Nvidia (NVDA) stock. Other models like Gemini are even more concentrated, holding a median of only 4.5 stocks.
AI recommendations appear to be primarily driven by the media attention that firms receive. Portfolio stocks receive nearly ten times as many news articles as the average Compustat firm, with a mean of 156,800 versus 16,500. Fractional logit regressions confirmed that the ability to grab attention within the universe of corporate news is a major driver of recommendations. Furthermore, portfolio stocks are dramatically larger than the full sample average, with a mean market equity of $289.3 million compared to $18.9 million. They also have substantially lower book-to-market ratios, with a mean of 0.24 compared to 0.71 in the full sample. Industry holdings deviate substantially from the S&P 500, with semiconductors (Chips) representing an average of 40.9% of AI portfolios.
In terms of performance, AI returns exceed the return of the S&P 500, but do not earn abnormal returns when considering trading costs or sector corrections. The six-month excess return over the S&P 500 was 1.748%, significant at the 5% level. However, this disappears once adjusted via DGTW to 0.644%, which is statistically insignificant. Active portfolios show non-trivial daily turnover, ranging from 3.13% for ChatGPT 5.0 to 9.05% for Claude. Implied holding periods range from 1.30 days for Claude to 5.08 days for Gemini. Among OpenAI models, ChatGPT 5.0 achieved the highest Sharpe ratio of 1.579.
Conclusions
The report concludes that understanding AI recommendations is critical as the general public relies on these platforms for major financial decisions. Preliminary results suggest that more oversight is needed to ensure people do not misuse this source of information and experience welfare losses. AI takes risk, recommends a narrow set of assets, focuses on specific industries, and does not appear to exhibit better performance than passive characteristic-based benchmarks. The assessment is ongoing, and the researchers plan to continue updating the data and revising conclusions based on new information and feedback from the academic and investment communities.