Research

A new NBER benchmark finds optimized agentic AI systems more than double the variation in stock returns they can explain around earnings announcements

Working Paper 35431 by Ralph Koijen and Bradford Levy proposes a real-time, out-of-sample benchmark that sidesteps look-ahead bias and market reflexivity, finding the best-optimized agentic AI systems push explained return variation from an R-squared of roughly 8% to close to 20% versus standard benchmarks, while also producing human-readable economic explanations for the moves.

· Source: NBER


The methodological contribution here matters more than the headline number, because asset-pricing research using AI has a well-known credibility problem that this benchmark is specifically built to dodge. Backtests on historical data are vulnerable to look-ahead bias, where a model appears to predict returns because it was tuned on data that already contains the outcome it is supposedly predicting, and market reflexivity, where a signal that worked historically stops working once enough capital trades on it, which is precisely the crowding dynamic covered in the alpha-decay research on this site today. Koijen and Levy's benchmark is real-time and out-of-sample by construction, evaluating how well a system explains stock returns around earnings announcements using only information that was actually available at the time, which is a meaningfully higher bar than most published AI-trading results clear.

Doubling the explained variation from roughly 8% to close to 20% sounds incremental until you sit with what an R-squared of 8% means for a professional trader: it means the vast majority of return variation around earnings announcements was previously unexplained by the benchmark models being compared against. Closing even a fraction of that gap with a system that also produces interpretable output, not just a better number but a stated economic mechanism for why the number moved, is a different kind of result than the marginal accuracy gains that dominate most machine-learning-in-finance papers.

The interpretability point is what separates this from a pure performance paper. The authors frame the benchmark as offering 'a data-driven method to identify missing variation factors and integrate empirical findings with economic theory,' which is a research agenda, not just a trading signal, the goal is using the AI system's output to find the economic variables that traditional asset-pricing models have been missing, then folding those variables back into theory that a human can audit and reason about, rather than treating the AI system itself as an opaque black-box edge to be deployed and protected.

Where this lands in the broader AI-in-markets conversation is as a counterweight to the alpha-decay argument covered elsewhere on this site today, that AI adoption compresses tradeable edges toward zero as more participants converge on the same signals. Koijen and Levy's benchmark does not directly contest that claim, a real-time, replicable methodology is exactly the kind of tool that accelerates the convergence the alpha-decay paper describes, but it does suggest the analytical value of agentic AI systems, in identifying genuine economic mechanisms rather than just short-lived statistical patterns, may persist even after any tradeable edge from the same models has decayed away. Watch whether other research groups adopt this real-time out-of-sample framework as a standard, since a shared, harder benchmark is usually what separates a durable research agenda from a one-paper claim.


Read the original: NBER - A new NBER benchmark finds optimized agentic AI systems more than double the variation in stock returns they can explain around earnings announcements. Commentary is the independent editorial view of Share Trading; the original article is credited to its publisher.