Continual learning trading models: what adapts, and how fast
Continual learning trading models adapt on different clocks. Finforge's Sophon-3 paper, July 2026, measured monthly rule re-search at 23% a year.
Continual learning trading models are systems that keep updating after they are deployed, and Finforge's Sophon-3 paper, Version 1.0, published in July 2026, makes a sharper claim than that. Its thesis is that the parts should not all update on the same clock. Table 2 of that paper lists four design choices, and the adaptation that measured largest was not in the deep learning layer at all. It was monthly re-search of the screening rules, worth 23% of active return a year in the paper's own pooled measurement.
Which is an awkward thing for an AI company to publish. The neural part is the part everyone sells. In this study it was the rule search that carried the measured edge.
What do continual learning trading models actually adapt?
The paper puts its thesis in one sentence. Markets are adaptive systems, so financial AI should be a causal, continually learning decision graph whose components adapt at different time scales while preserving point-in-time correctness.
Three terms carry the weight there. Causal means a step may only use information that existed when the decision was made. Point-in-time correctness means the calculation runs on the data as it stood on that date, not on the restated, survivorship-cleaned version sitting in a vendor database today. Different time scales means the components are allowed to move at different speeds, and are expected to. Anyone who has run a research desk knows which of the three costs the most to get right.
Which clock paid, and which one did not?
The slow clock sits in the Screener, the step that decides which names are even eligible. Table 2 calls the mechanism screener-first adaptation: a rolling explore and exploit search over interpretable rule families. Explore and exploit means the system keeps trialling new rules while continuing to use the ones already working. Interpretable means a human can read the rule, not just run it. So selection itself keeps learning, month after month.
That monthly re-search is recorded as the largest measured adaptive contribution in the paper, at 23% of active return a year against a simulated copy of the same agents with their rules fixed once at inception. Active return is portfolio return minus the return of the benchmark universe the agent trades. Two other removals bracket it: taking the Screener out entirely cost 45% a year, and forcing equal-weight position sizing instead of sizing by forecast magnitude cost 36% a year.
Now the fast clocks. Retraining the forecaster continually scored +4% of active return a year with a confidence interval that spans zero, so the paper does not claim it. Adaptive ensemble weighting came out as a precise null. Finforge built both mechanisms and then printed both as unproven at this sample size.
Those figures are pooled across agents and come from simulated reruns of the same system with one component removed, over the same sample and against the same benchmark universes. They are model comparisons, not an account statement. Past performance is not a guide to future returns.
Why do backtested trading strategies fail live, and what guards against it?
Two of the four rows in Table 2 exist for that reason. Walk-forward forecasting means repeated train and predict chunks with strict validation: fit on one slice of history, predict the next, roll the window, repeat. The paper's justification is blunt, that it treats changing regimes as the default case rather than as an exception to patch later. The second guard is what Finforge calls no-future-bleed ensembles, where the blending weights are resolved strictly before the current period's losses are observed. That is there to prevent same-period loss leakage, which is the polite name for a model being handed credit after the answer is already known. We covered the wider question separately in three reasons backtested trading strategies fail live.
Why a graph instead of a pipeline?
Because different clocks need different wiring. Table 2's first row is a composable node graph: causal operators that can be assembled into any DAG, short for directed acyclic graph, a flowchart where work moves forward and never loops back. Chains, fan-in and fan-out are all allowed, and the familiar four-node chain is described as one instantiation of it. The stated benefit is operational rather than clever, that many workflows can be expressed without rewriting the system.
Questions people ask
Is continual learning just retraining the model every month? Not in this paper. Continual retraining of the forecaster scored +4% of active return a year with a confidence interval spanning zero, while the monthly rule re-search inside the Screener measured 23% a year. Same idea, two different layers, two very different results.
If the model keeps changing, how is it tested? Through walk-forward forecasting with strict validation, where each prediction is made on a chunk the model has not seen, and through point-in-time correctness, where every decision is scored on the data that was visible that day.
Does adaptation remove market risk? No. It addresses one failure mode, a model fitted to a market that has since moved. Everything else stays exactly where it was.
Finforge Research wrote the Sophon-3 paper and runs its trading agents in public, wins and losses included. Four Sophon agents have traded since April 21, 2026, on founder capital in Alpaca paper accounts, so no customer money is being traded before launch. You can read the live results, benchmarks and drawdowns for every Sophon agent, and the Sophon-3 research summary holds the full Table 2.
The paper is public. So are the agents.