Educational

AI trading strategies: how to spot a survivorship-biased backtest

AI trading strategies can cheat with survivorship bias. Finforge's Sophon-3 paper, July 2026, restricts picks to point-in-time index members to stop it.

AI trading strategies can lie, and the quietest way is survivorship bias. Judging a system only on the names that survived. Finforge's Sophon-3 paper, Version 1.0, published in July 2026, guards against it with a data contract that refuses to show a record before its own timestamp says it existed. One written rule, and a chart that looks unkillable stops being enough.

Survivorship bias is the gap between the market that actually ran and the tidy list you'd see today. Plenty of backtest databases hold only current index members. The companies that fell out, got bought, or went under just vanish from the record. A strategy that always happened to own the winners looks brilliant. It never shows you it would've also held the dead ones.

What does an AI trading strategy get to see, and when?

The Sophon-3 paper calls its answer the data contract. Every node reads a timestamped contract, and a record enters the model only if its observable timestamp is at or before the moment of decision. A fundamental enters at its filing or release time. News enters at publication. A price bar enters only after it closes. Timestamp semantics are part of the procedure, written as Π, and stored with each run manifest.

That closes the window for lookahead, the polite word for a model that used information it couldn't have had yet. A price still moving is not allowed in. News not yet published is not allowed in. The contract is a gatekeeper, and it's built into the design, not bolted on later.

Why point-in-time index membership stops a survivorship-biased backtest

The paper pairs the data contract with a universe rule. At any moment t, the assets an AI trading strategy may pick are restricted to the constituents of the index as it stood at t. If a stock wasn't a member on the day of the decision, the strategy can't select it. The paper says this materially distinguishes Sophon-3 from survivorship-biased backtests.

Put plainly, the universe is dated the same way the data is. The model works on the index as it looked that day, not the index as it looks now with the casualties edited out. Point-in-time means exactly that. The picture as it stood, not the tidy version corrected by hindsight.

Prices: split-adjusted for the model, familiar on screen

Even the prices get the same treatment. The warehouse stores raw and split-adjusted fields. The model-facing features use forward split-adjusted prices, so a split doesn't show up as a sudden cliff in the series. The display maps back to the familiar prices you'd actually recognise. Two views of the same number, one for the machine and one for you.

Why the benchmark matters as much as the strategy

A clever AI trading strategy can still flatter itself at the scoring stage. The paper calls the benchmark convention the most common source of illusory active return in long-only equity research. Its fix is symmetry. Portfolio and benchmark get the same dividend treatment inside each evaluation domain.

The simulated backtests don't credit dividends to the portfolio, so they're compared to the price index, ex-dividend on both sides. The live brokerage account does receive dividends, so its net asset value is total-return and it's compared to the total-return index. Symmetry is the requirement, not one particular index. On the growth-tilted books evaluated in the paper, whose yield runs at or below the Nasdaq-100's roughly 0.7 to 0.9 percent a year, a price-basis figure overstates active return by at most that gap. Ablation deltas, the tests that remove one component, come out the same on any basis.

Questions people ask

What is survivorship bias in an AI trading strategy? It's judging a strategy on a universe that only remembers the stocks that survived. A backtest that quietly edits out the losers overstates how good the strategy really was. Point-in-time index membership is the fix. The model only trades names that were actually in the index on the day.

What does point-in-time mean in plain words? Every decision is scored on the data as it stood on that date, not on the revised, hindsight-corrected version sitting in a database today. The data contract and the universe rule both enforce it.

Why does the benchmark get its own section in the paper? Because active return only means something relative to a benchmark measured on the same dividend basis. If the strategy keeps dividends and the index doesn't, or the reverse, the comparison flatters one side for free.

Finforge runs four Sophon agents in public and publishes what they do, wins and losses. They have traded since April 21, 2026, on founder capital in Alpaca paper accounts, so no customer money is being traded before launch. You can check the live results, benchmarks and drawdowns for every Sophon agent, and the Sophon-3 research summary sets out the full data and universe protocol. Past performance is not a guide to future returns.

Published as a research paper. Coming to your phone.