Research explainers

Does backtesting work? Sophon-3 lists what it can't prove

Does backtesting work? Finforge's Sophon-3 paper, Version 1.0 July 2026, grades two of its own mechanisms unproven and prints the numbers.

Does backtesting work? Up to a point, and the Sophon-3 paper from Finforge Research, Version 1.0 published in July 2026, shows exactly where that point sits by printing a table of its own claims with two of them graded Not yet proven and Null so far.

Most quant research buries the weak results. This one puts them in Table 1, next to the strong ones, with a column explaining why each is still open.

Does backtesting work if the market has already moved on?

The paper opens by saying the central problem in systematic investing is not prediction. It is prediction while the mechanism generating the data changes, while you can see only part of the picture, while your own trades feed back into the market, while operational constraints bite, and while the evidence about whether a signal was useful arrives late. Five problems at once. A static backtest answers none of them.

The failure list that follows is ordinary, which is the point. The regime changes. The eligible universe drifts, meaning the list of names you were allowed to buy is no longer the list you tested on. Costs rise. Live execution diverges from research.

What does Sophon-3 offer instead of a better backtest?

Contribution one is a formal definition of leakage-free continual learning for a decision graph, written in terms of a filtration, the mathematical way of saying a decision may use only what was knowable at that moment. The twist is in how evaluation is defined. You freeze the procedure, not the model state. The models keep learning. The rules about how they learn are locked.

Contribution two is a forecast ensemble whose weights depend only on past realized losses. Finforge calls that no-future-bleed. If a blend of models looks clever because it quietly knew which model would win later, the weights were contaminated and the result is fiction.

Contribution three is the one a professional will care about most. Results are dated to an architecture freeze, separated by where the evidence came from, and scored at the final tradable-portfolio level rather than at signal level. Signal accuracy is easy to dress up. A portfolio you could actually have held is not. The paper also includes a real-money out-of-sample window.

Which parts of the system survived being removed?

Finforge tested by subtraction. An ablation removes or freezes one component of an agent, where an agent is one deployed copy of the Sophon-3 graph trading a single benchmark universe. Effects are paired daily differences of active return, which is portfolio return minus benchmark return, taken as the agent minus the stripped variant. Pooled means the three agents' difference series are averaged before any statistics are run.

Remove the Screener and the pooled cost is 45% of active return a year, significant in all three agents. Force equal-weight sizing instead of sizing positions by forecast magnitude and the pooled cost is 36% a year, again significant in all three. Let the Screener keep re-searching its rules rather than fixing them once at inception and the pooled gain is 23% a year.

Two mechanisms did not clear the bar. Continual retraining of the forecaster came in with a 95% confidence interval running from -5% to +13% a year, which spans zero, so the paper grades it Not yet proven. Adaptive ensemble weighting came in at -0.0% a year, give or take 0.6%, graded Null so far and kept anyway as insurance against a change of regime. Finforge built both. Finforge published both as unproven.

Those figures compare each agent against a simulated rerun of itself with one part removed, over the same sample and against the same benchmark universe. They are model comparisons, not an account statement. Past performance is not a guide to future returns.

What has the paper not established yet?

Table 1 has a fourth column, and it is the useful one. External reproducibility of the leakage-safe graph rests on a companion parity paper. The full search-space multiplicity ledger, the accounting of how many variants were tried before one was kept, sits in a companion selection paper. The sizing result has no isolated mechanism yet and may be proxying volatility or dispersion. The continual Screener finding has only a short design-frozen confirmation behind it. A leak diagnostic for the ensemble weighting is still pending.

The scope section is blunter still. The paper does not assert that any model beats all markets in all regimes, that positive backtests prove future profitability, or that continual learning eliminates market risk. It calls its own live track record explicitly short and not yet statistically conclusive.

The motivating hypothesis is stated as testable rather than universal: that causal continual-learning graphs are a more appropriate architecture for markets that keep moving than once-trained models or fixed rules. Finforge anchors it in the Adaptive Markets Hypothesis and in concept-drift research, where the link between inputs and the thing you are predicting changes over time.

Does a positive backtest prove a strategy will make money?

No, and Sophon-3 says so in its own words. A backtest describes how a procedure behaved on a sample it has already seen. What a dated freeze adds is a cleaner question, because everything recorded afterwards ran against a design nobody could edit once the results were visible.

What is an ablation, in plain terms?

You take a working system, switch off one piece, rerun the same period, and measure the gap. It is how you find out whether a component earns its place or just makes the diagram look busy. In Sophon-3 the Screener earned its place by a wide margin. The ensemble weighting has not, so far.

Who is behind these numbers?

Hans Dahlstrom, Adam Tittenberger and Ted Bjorling at Finforge Research, in a paper titled Sophon-3: Continual-Learning Deep Models for Non-Stationary Financial Markets, Version 1.0, July 2026. Finforge runs its agents in public and publishes what they do, wins and losses. The four Sophon agents trade founder capital in Alpaca paper accounts, so no customer money is being traded before launch. You can check the live results, benchmarks and drawdowns for every Sophon agent, and the Sophon-3 research summary carries the tables.

Published as a research paper. Coming to your phone.