AI trading strategies: how to tell what's really working
AI trading strategies: Finforge's Sophon-3 paper, July 2026, runs ablations to tell which part earns the edge, and prints what it cannot prove.
AI trading strategies can be climbing and you still cannot say which part of the system is earning the gains. Finforge's Sophon-3 paper, Version 1.0, published in July 2026, opens its evaluation section with exactly that warning. A live record shows a system operates. It does not show why. To find that out, the paper strips the strategy apart.
AI trading strategies: what gets the credit?
The trick is called an ablation. Take a working agent, remove one piece, rerun the same stretch of history, and measure the gap. In the paper, every ablation is a paired daily difference: the agent minus its own stripped version, over the same dates. Never two independent guesses. Same sample, same benchmark, only one part changed.
The claim that carries the science is the ablation suite, not the live curve. That sentence is deliberate. Finforge runs four agents live, split by risk class into high, medium and low. Three of them, Sophon Apex, Sophon Core and Sophon Surge, hold a forecast step and form the test panel. Each panel agent is a full walk-forward run from early 2023 to 1 July 2026.
Where the honesty checks sit
The paper prints two warnings about its own numbers. First, with 21 paired tests at the 95% level, one or two isolated hits are expected by chance. A single strong cell proves nothing. The evidence is consistency across all three flows. Second, the three flows' active returns correlate at an average of 0.31, so they behave more like 1.9 independent flows than three. Finforge refuses to count them as three separate proofs.
There is also a short post-freeze check, from 21 April to 1 July 2026, about 49 trading days. The paper flags it as a directional check only. A stump of a quarter is not a verdict.
The verdict, in three tiers
Results land in three tiers. Demonstrated means the merged confidence interval rules out zero. The screener's contribution came in at about 45% a year of active return, sizing positions by forecast magnitude at about 36% a year, and the monthly re-search of the screener's rules at about 23% a year.
Inconclusive means the interval crosses zero. Continual retraining of the forecaster measured +4% a year with a 95% interval from -5% to +13%, so it could be helping or not. Finforge grades its own flagship habit as not yet proven. The precise null is sharper. Adaptive ensemble weighting measured -0.0% a year, give or take 0.6%, and the paper keeps it anyway, as cheap insurance against a change of regime.
Active return is the portfolio's return minus the benchmark's return. These ablation figures are model comparisons from the paper's re-runs, not an account statement. Past performance is not a guide to future returns.
Why this matters for judging any AI trading strategy
The habit to copy is the disclosure. A strategy that shows its live wins but hides the reasons, and hides the tests that failed, is doing the opposite of this paper. Finforge names the number it cannot prove and says so in plain words.
Finforge runs four Sophon agents in public and publishes what they do, wins and losses. They have traded since 21 April 2026 on founder capital in Alpaca paper accounts, so no customer money is being traded before launch. You can check the live results, benchmarks and drawdowns for every Sophon agent, and the Sophon-3 research summary carries the full evaluation protocol.
Questions people ask
How do you know an AI trading strategy is working?
A live record shows it runs. To see which part earns the edge, run ablations: remove one component, rerun the same period, and compare day by day. Finforge's Sophon-3 paper grades each result as demonstrated, inconclusive, or a precise null.
What is active return in an AI trading strategy?
Active return is the portfolio's return minus the benchmark's return. In the paper every ablation is a paired daily difference, agent minus stripped variant, over the same dates.
Why publish results you cannot prove?
Because the reader sees the line between what is demonstrated and what is a hunch. Continual retraining measured +4% a year with an interval that spans zero, so it stays unproven. Showing that line is the trust.
Published as a research paper. Coming to your phone.