Educational

AI trading strategies: why a short record can't be proven yet

AI trading strategies: over 41 trading days even an IR of 3 gives a t-stat near 1.2, so a short live record is not proof on its own.

AI trading strategies can post a clean three-month win streak and still have no statistical proof that they work. Finforge's Sophon-3 paper, full text (v1.0, July 2026), lays out that math in its section on statistical inference. Over about 41 trading days, roughly 0.16 years, even an information ratio of 3, an elite grade, comes out to a t-statistic near 1.2. The paper does not present the live record as significant on its own.

AI trading strategies: why a short live record falls short of proof

A t-statistic is a rough signal-to-noise score for a result. The higher it is, the less likely the result is random. The paper gives the link between the live window and the score: the t-statistic on an information ratio is about IR times the square root of the years. Shorten the years and the score collapses, no matter how good the strategy looks. Even IR equals 3, a grade most funds never reach, only clears a t of about 1.2 over 0.16 years. That sits below the bar most researchers would call significant, so the paper refuses to call the short live record proved.

Why a smooth winning curve is not many confirmations

The second trap is counting overlapping data twice. A rolling information ratio sampled daily over 126 bars is not 126 independent checks. It is heavily autocorrelated, meaning each day's value carries most of the previous day's. So a smooth multi-month curve is roughly one independent block of evidence, not a stack of confirmations. The paper presents it as evidence of stability, not as repeated significance.

Why four agents beating the benchmark count as closer to two

Now the counting gets subtle. "Four agents beat the benchmark" is four pieces of evidence only if their active returns barely move together. Active return is the portfolio's return minus the benchmark's return, and it measures the edge a strategy earns. To find how many real bets the panel holds, the paper uses the effective number of independent bets, Neff, roughly N divided by 1 plus N minus 1 times the average correlation, rho. For the three-flow ablation panel the realized numbers are an average correlation of 0.31 and an Neff near 1.9. Four names that move together are closer to two independent opinions. The paper reports the correlation matrix so it cannot over-count "all four".

How the paper stops luck from looking like skill

Multiple testing is the last honesty check. Run enough comparisons and a couple of winners are expected by chance alone. Here the ablation suite runs 21 paired tests at the 95% level, which implies one or two chance significances. So the paper rests its conclusions on cross-flow consistency and pooled confidence intervals, not on any single strong cell. Add an adaptive architecture that inflates the researcher's degrees of freedom, and the caution is deliberate. The paper says exactly that: it is deliberate about power.

Finforge runs four Sophon agents in public and publishes what they do, wins and losses, since April 21, 2026 on founder capital in Alpaca paper accounts. No customer money is traded before launch. You can check the live results, benchmarks and drawdowns for every Sophon agent, and the Sophon-3 research summary sets out the evidence standards in full. Past performance is not a guide to future returns.

Questions people ask

Does a short winning record prove an AI trading strategy works?

Not on its own. Over about 41 trading days, roughly 0.16 years, even an information ratio of 3 yields a t-statistic near 1.2, below the usual bar for significance. The paper treats the live record as evidence of stability, not proof.

Why do four agents beating the benchmark not count as four proofs?

Because their active returns move together. At an average correlation of 0.31, the effective number of independent bets for the three-flow panel is near 1.9. Correlated results are one opinion spoken several times.

What is a t-statistic?

A rough signal-to-noise score. The higher the number, the less likely a result is random. The paper notes the t-statistic on an information ratio is roughly IR times the square root of the years, which is why a short window stays weak.

Published as a research paper. Coming to your phone.