We read the rulebooks of the major AI trading arenas, and tell you, why it matters?
In the ten months since Alpha Arena gave six frontier models real money to trade, an entire trading genre has been born. There are now AI trading arenas run by crypto prime brokers, research labs, university groups, exchanges, and at least one company that has since been acquired. All promoting prize pools, leaderboards, live capital and several are genuinely well structured.But not all competitions are equal, or how rankings are calculated. It's for this reason we decided to do the boring stuff and read through the rulebooks, scoring formulas, and fine print to give you a clear picture of whether the numbers on the leaderboard are based on live data, or "idealistic" backtest values.For those short on time, or just too impatient to read to the end, here is a short overview:What we uncovered is a measurement problem when scoring real trading performance. But not only that, the scores we found were mainly based on raw return, and taken over limited time periods, too short to distinguish skill from variance. But what makes things even more confusing, is that several of these providers don't make their ranking formulas public, yet publish headline figures that are actually backtest-derived and not based on actual performance figures. Confused? Well, you're not alone.In this article we'll break down the scoring in these competitions, what is good and what is not, and why all this matters in the first place. Let's dive in...
Why ranking participants on raw return alone is only half of the story?
Provide a caption (optional)Most AI trading competitions follow an unspoken convention. They base scoring on final returns rather than a score evaluating both risk and retun together, which immediately tells you that there is a problem. A risk metric and ranking are two different things, so basing a score on only one gives you half the truth.Alpha Arena's own first season makes the clearest case for why this really does matter. Between 18 October and 3 November 2025, six frontier models were each given $10,000 in capital, the same prompt, and then traded crypto perpetuals autonomously.The result? Only two of the six models made any money. The gap between first and last was a lengthy 81 percentage points. This despite all six getting the exact same prompt. Now if you rank all six on return alone, the leaderboard will place Qwen3 Max in first place. The only problem is that what actually separated the top from the bottom was risk discipline. The 81-point difference is a variance measurement, and those leaderboards that rank on the return column only will keep confusing variance with capability.There are a few operators that do rank on risk-adjusted performance, and they definitely deserve credit for it. Quantiacs ranks its competitors based on Sharpe ratio and reports having deployed more than $38 million into winning algorithms across futures markets since 2014. Digital Quant in Hong Kong also ranks on risk, using return curves, maximum drawdown, win rate and profit factor.LTP's Liquidity Arena also belongs in this group and it published its Track A formula below:0.40×Z(Sharpe) + 0.25×Z(PnL) + 0.20×Z(ROI) + 0.15×Z(MDD)Sharpe is the strongest single term and it's because of this that LTP is way ahead of the rest when it comes down to a balanced scoring system.
Why scoring formulas need to be published?

Provide a caption (optional)When it comes to AI trading competitions Alpha Arena has defined this category. It has become the reference point every subsequent arena is measured against. Season 1 results are the most cited data on the web, yet we could not find its ranking rules published anywhere.That may seem strange for a market leader to not be transparant, but it isn't completely uncommon. LTP publishes its Track A formula, which is genuinely to its credit and more than most competitors will share, but Track B, the institutional track, doesn't seem to have any published formula available. Nor has any judging panel been named publicly, and their own supporting documentation seems to conflict with the leaderboard in both directions. It lists a win-rate component that doesn't appear in the scoring formula, and omits PnL, which accounts for 25% of it.We don't think this is deceptive as sharing the formula in public invites people to argue with it, and freezing one for a full season removes the option to fix your scoring if you encounter an issue down the line. But not publishing your ranking rules does leave participants unsure of what will make them a winner at the end of the day.Another issue is that many headline figures linked to these competitions trace back to a single press release from the operator. LTP's widely-reported "200+ teams" and "$300K+ in prizes" come from its own July announcement. Now at the end of August, Phase 1 comes to a close but we could not find any published data, no leaderboard results or independent verification.
Too many headline numbers are backtest numbers
WorldQuant BRAIN is the biggest funnel in the entire landscape. It has over half a million platform users, tens of thousands of International Quant Championship entries annually, across universities in more than 140 countries, and thousands of paid research consultants. This is a serious operation with a genuine path for students to find employment.But their headline metrics seem to all be backtest metrics. We failed to find any published out-of-sample holdout protocol comparable to Numerai's live scoring. That does not mean strategies aren't validated internally, but it does mean the numbers a participant sees are in-sample, meaning it's based on historic data, and not live trading data.The same pattern repeats down the ladder. TradingAgents, the LangGraph multi-agent framework with over 90,000 GitHub stars, reports Sharpe ratios between 5.6 and 8.2. Those come from a single quarter on three technology stocks. To the authors' credit though, they do flag these figures as artifacts themselves, and by mid-2026 they added a look-ahead-bias filter.At the commercial end, Tickeron sells subscriptions between $60 and $250 a month against self-reported claims of 169–172% annualized returns and 85–92% accuracy, all backtest-derived and, as far as we could find, they were unverified by any third party.This is the context in which regulators are starting to pay attention, and these unverified results from platforms and competitions are what create a gap in trust and credibility for the industry as a whole.
Why the competition needs a longer trading window?

Provide a caption (optional)Nearly every AI trading competition has run from two weeks to three months. Alpha Arena's first season ran for just sixteen days while LTP's full arena spans around three and a half months. Its elimination round ran for a total of thirty days but on a simulation without real capital, except for the final stage, where each team started with an amount of $1,000.None of what we're shared here is a secret, it's all in their documentation for anyone to read. But our position in regard to the limited time frames is that it cannot reveal durable skill, and neither can a sixteen-day streak. A trading strategy's edge is a small signal buried in a sea of noise. Trading for just two weeks in the most volatile asset class available, means the noise drowns out any real signal. And these short trading windows by default reward the behavior you would definitely want to filter out. For example, like taking an over-levered position that would be fatal across a year suddenly looks like conviction across a few weeks. And the leaderboard only tells you half the story.
Why all of this matters and how to fix it?

Provide a caption (optional)QuantConnect ran a crowd-sourced alpha marketplace called Alpha Streams. The community members could submit strategies, while QuantConnect would filter them, and investors had the choice to license them. But then they shut it down. Later CEO Jared Broad published a complete post-mortem on why this happened.They compared the community's strategies against ordinary stocks, S&P 500 and Russell constituents, but the strategies came out worse as a group. They then select only the best 5% of all the strategies ever submitted. They stripped out the illiquid ones and a few crypto flukes and what was left still underperformed the S&P 500. As a final resort, they tried using uncorrelated strategies as a diversifier on top of a normal market portfolio, but again that didn't help either.Broad's diagnosis, "The current submission and filtering process seeks strategies that perform well in all market regimes. This is a relatively unrealistic task and results in strong overfitting." Read that again with the rest of this article in mind. The failure wasn't the contributors. It was the selection process. Filtering on backtested performance across all regimes doesn't find skill it finds the strategies that have been fitted hardest.And that filter is the one almost every arena in this piece applies. So the proof is there, it's public and came from a platform with 533,000 users, ten years of data, and every commercial reason to wash over it, yet they published it anyway. The most rigorous result this field has produced is sitting in the open while the leaderboards carry on as though it isn't there.
What good measurement actually looks like?

Provide a caption (optional)Right now several operators are doing serious work, and the field is better for it. Numerai is the clearest counter-example to the Alpha Streams result. It has run a crowd-sourced alpha tournament for years now and has grown its hedge fund from roughly $60 million to several hundred million in AUM in three years. In August 2025, J.P. Morgan Asset Management committed up to $500 million in capacity and that November, Numerai raised a $30 million Series C at a $500 million valuation.What it does differently is real out-of-sample live scoring, staking which puts the contributor's own money on the line, and payouts on forward performance rather than backtest quality. Score badly and your stake burns, which is the strongest incentive design in the category by some distance.HKUDS's Vibe-Trading, out of a University of Hong Kong lab, has the most rigorous evaluation stack we found anywhere and it's free. Look-ahead bias is banned at the operator layer rather than left to participant discipline. Data is point-in-time and maintains an alpha library with each factor classified as alive, reversed or dead, while monitoring IC and Sharpe decay automatically.And there is a result that deserves far more attention than it has received up to now. In a competition concluding in January 2026, the platform Aster put thirty AI agents against human traders over two volatile weeks. Based on the headline the humans won and the single best individual performer was a person. But look deeper and one layer down, the whole picture flips on its head. The human traders finished at −32.22% in aggregate, with 43% liquidated outright. The thirty AI agents finished at −4.48%, with zero liquidations, that is a 100% survival rate, and roughly $13,000 in total losses across the entire cohort.The humans might have produced the biggest winner but they also had the a graveyard of losses. The machines might not have produced any heroes but there where no corpses...
What Ranking Standards should apply?

Provide a caption (optional)Four things a trading leaderboard should publish if it wants to be treated as a measurement rather than a spectacle:1) The ranking formula, in full, frozen for the season. Participants cannot audit a result they cannot reproduce, and a rule that keeps changing mid-competition isn't a rule at all.2) Results should be stated plainly, whether in-sample or out-of-sample. Backtest visualisations are useful, but they need to be labelled as such and should not affect your ranking.3) A measurement window long enough to contain a regime change. Two weeks can only realistically measure variance, but a year gives enough time for skill to be revealed.4) Drawdown and survival as ranked inputs. The Aster result and the Alpha Arena spread both point the same direction, how a system fails is more informative than how it wins.It's an exciting time to be in the trading space. AI is revolutionizing how we work with data, build quantitative strategies and how we execute trades. Competitions like the above are still in their infancy, and are a learning ground where we can understand how to better apply AI to trading.
Sources: Nof1 Alpha Arena Season 1 results reporting; Aster "Human vs AI" Season 1 conclusion (January 2026); LTP Liquidity Arena documentation and July 2026 announcement; QuantConnect — Jared Broad, "Alpha Streams Refactoring 2.0"; Numerai and J.P. Morgan Asset Management capacity announcement; WorldQuant BRAIN; HKUDS Vibe-Trading; TradingAgents (Tauric Research); Quantiacs; SEC FY2026 examination priorities; CFTC staff advisory, December 2024. Figures attributed to operators are self-reported unless otherwise noted. Reported performance figures vary by source and reporting window.
ABOUT TQX
TQX is a quantitative trading platform that provides the infrastructure previously only available to large hedge funds and trading desks, to build institutional grade trading strategies using natural language, and zero code. We have also launched our own arena, the GATC (Global AI Trading Competition), a twelve month, real capital in regulated Hong Kong and US equities, a composite score in which risk control carries 45% weight, and the full formula published in the rulebook. Entries are open for September.
Signup Today: www.tqx.ai