Research
Ten leading AI models were given identical starting capital and set loose on live foreign exchange markets, and the results split further apart than expected
HKU Business School's Artificial Intelligence Evaluation Lab connected 10 leading LLMs, including models from OpenAI, Anthropic, Google, Alibaba, Moonshot AI, Zhipu AI, ByteDance and DeepSeek, to live FX data through its Agentic Trader system, with identical starting capital in an identical market environment, and found Alibaba's Qwen produced the strongest profit.
What makes this study genuinely useful, rather than another leaderboard exercise, is the controls. Ten models from eight different labs, same starting capital, same live market, same window of time. That removes the usual excuse for divergent AI trading results, that one system had a data advantage or a friendlier backtest period. When conditions are held constant and performance still spreads this widely, the gap is coming from the models themselves, specifically how they reason under continuously changing conditions rather than how they perform on a static benchmark.
Qwen winning is the headline, but the more interesting finding sits underneath it: performance gaps between models widen as the task demands continuous interpretation and real-time strategy adjustment. That is a meaningfully different claim to saying one model is simply smarter than another. It suggests the skill that separates a good AI trading agent from a mediocre one is less about raw market knowledge and more about a specific kind of adaptive reasoning, updating a live position as the environment shifts rather than reasoning once and holding a static view.
The lab running this, HKU's AI Evaluation Lab under Professor Jack Jiang, sits inside a wider Centre for AI, Management and Organization with a research programme specifically built around AI agents in business and economic settings. That institutional framing matters for how seriously to take the result. This is not a single paper or a vendor benchmark designed to flatter one model, it is a purpose-built evaluation environment intended to run this comparison repeatedly as models improve, which means today's Qwen result is a snapshot, not a verdict.
The unresolved question is whether a live FX result generalises to equities, where liquidity, news flow and correlation structures behave differently, and whether the model that wins this round would still win once every serious lab starts training specifically to beat exactly this kind of evaluation. A benchmark that becomes well known enough starts shaping the thing it measures, and Agentic Trader is exactly detailed and public enough to become a target rather than just an observer.
Read the original: HKU Business School - Ten leading AI models were given identical starting capital and set loose on live foreign exchange markets, and the results split further apart than expected. Commentary is the independent editorial view of Share Trading; the original article is credited to its publisher.