Research blog

Predicting How Millions of People Change Their Minds

Closing the outer loop of self-improvement

Ordinary self-improvement happens inside change we control. The outer loop happens inside change we do not. This post puts a model in the second kind, and asks whether it gets better at predicting how a crowd of millions will move.

hours we made a call on
6,984
outcomes we learned from
4,839
days of dated snapshots
390
update rounds completed
12

Problems we invent, facts that arrive

The inner loop is fast because we arrange it. The task, the grader, the next item, even how the environment moves — all of that is ours. Code compiles or it does not. A proof checks out. Self-play closes in milliseconds. Models improve fastest there because nothing outside the run is allowed to rewrite the problem.

The outer loop is slower because the outside can move on its own, in ways that were never on the syllabus. A headline lands. A crowd reprices. A contract settles. Nothing about that sequence was written to teach the model. Call the world's answer a verdict: what actually happened next, arriving on the world's clock rather than ours. Settlement is the slowest and most final kind, but it is not the only one. An hour after a call, the crowd has already moved or failed to move, and that movement is a fact nobody in the lab authored.

ONE CONTRACT Will the central bank cut rates by September? 62¢ YES · 62¢ NO · 38¢ the price is the crowd's probability ≈ 62% THEN A FACT ARRIVES inflation above consensus … and it reprices many of them at once Rate cut by September 62% → 41% ▼ 21 pp Recession within the year 18% → 26% ▲ 8 pp Equity index at a new high 55% → 48% ▼ 7 pp Incumbents hold the legislature 51% → 47% ▼ 4 pp knock-on — moved by the row above, not by the news pp = percentage points
Figure 1 A contract is a belief with a price — and beliefs move together. Illustrative magnitudes. Three of these reprice on the news directly. The fourth, drawn dashed, moves only because the row above it did — the legislature reprices because the growth outlook just changed. The first three are a forecasting problem. The fourth is why this is infrastructure and not just a model: nothing scores it correctly unless the link between the two contracts was recorded at the time. This run scores each contract on its own.

A crowd changing its mind is one such change, already carrying a price. We ran a 7B model through that loop for twelve months of recorded calls. What follows is the instrument reading — what the loop actually bought, and where it did not pay.

What the loop bought

One row is one contract at one hour, together with the thirty unfiltered wire headlines that preceded it. The model answers with the name of a bucket — one of nine bins for how far the next move lands above or below a frozen baseline that reads the current price and nothing else. Every prompted comparison below is scored on the same 2,145 held-out rows against that baseline. A model whose answer cannot be parsed is scored as the baseline, which is zero by construction.

A

Prediction error against the baseline

Our 2,145 held-out rows · Share of baseline squared error removed

  • claude-haiku-4.5
    −0.69%
  • deepseek-v4-pro
    −3.69%
  • gpt-5.6-sol
    −6.68%
  • claude-opus-5
    −27.44%off scale
  • ours · 7B, self-distilled
    +2.51%
← worse than doing nothingbetter →
Prompt-only comparisons: no tools, retrieval or memory. Zero = the price-only baseline.
B

How much the predictions vary.

Standard deviation of predicted and observed price changes · Probability units

  • Observed change
    0.104
  • ours
    0.0214
  • gpt-5.6-sol
    0.0557
0.01 = one percentage point. The outlined bar is the observed-change reference.
Figure 2 Our held-out results. Panel A compares prompt-only models and our self-distilled 7B model on the same rows against our frozen baseline. Panel B compares the standard deviation of predicted changes with the observed-change reference.

On our held-out rows, every prompt-only comparison finished below zero. Our self-distilled 7B model finished above it: +2.51%.

The numbers are small because the target sits on the noise floor. Most of what a crowd does in the next hour is drift back toward where it already was. Take out the drift and what is left is small, and the only part that carries information. On these rows a crowd revises by about ten points in an hour; the drift-only baseline already accounts for nearly all of it. The 2.51% is the share of what that baseline still gets wrong that reading the news takes back — the part of the crowd's change of mind that was in the headlines and not in the price.

The models that lose are not the ones that miss the direction. gpt-5.6-sol ranks outcomes better than ours does, picks the right bucket more often, and calls direction more often — it beats us on every measure of understanding we have — and it lands nine percentage points below us on this scale. What separates them is how hard they commit. It moves two and a half times as far as we do on a target that is mostly noise, so it is right about where the crowd is going and wrong about how far, and squared error charges it for the second part. claude-opus-5 makes the same point at the other extreme: it picks the correct bucket more often than anyone in the table, and has the worst error in the table. Picking the bucket and predicting the move are not the same skill.

Our model wins by knowing when not to speak. On four rows in five it answers that the news changes nothing. That is not an abstention but a judgment of no effect. The entire margin comes from the fifth row: the hours where the crowd was about to do something the price alone did not imply, and where it lands closer than the baseline 55% of the time. That is a smaller claim than forecasting, and a more useful one: it is a filter that has learned which arriving facts are worth acting on.

one row where it did speak · equity-index contract · 2026-02-13 14:00
+0.048the baseline expected
−0.037the model said
−0.060it actually moved
0.023model error, against the baseline's 0.108
… the recent price history shows a significant drop from 0.570 to 0.525 over the last few hours … the relevant headlines provided are: “U.S. January Unadjusted CPI 325.252, forecast 325.408, previous 324.054” — “U.S. January Seasonally Adjusted CPI MoM 0.2%, forecast 0.30%, previous 0.30%” …

Its own words, from the run. Out of thirty unfiltered wire items it reaches for the two inflation prints and commits to a move against the baseline, which had the contract rising about five points. It fell six. This is what the fifth row looks like when the model decides to speak.

A separate study, FutureSim, examines agents with search, file memory, and the freedom to decide when to revise their forecasts. The agents work hard at it: hundreds to thousands of tool calls and more than ten million tokens each. Search helps. Memory helps. More reasoning effort helps. In the five-agent experiment shown below, three finish below the score for declining to predict, and two finish slightly above it. In a separate test that starts every agent from the same deliberately weak forecast, none climbs back to zero.

External evidence · FutureSim

330 forecasting questions · Brier skill score · Search, memory and self-chosen update timing

  • GPT 5.5
    +0.05
  • Opus 4.6
    +0.02
  • GLM 5.1
    −0.01
  • Deepseek V4 Pro
    −0.02
  • Qwen 3.6 Plus
    −0.07
← worse than doing nothingbetter →

Zero = abstaining from prediction.

FutureSim · Five frontier agents in their native harnesses, final day · 500–4,000 tool calls and 10M+ tokens each.

The single number does not mean what it looks like. Split those same held-out rows twelve ways, by the month they came from, and the aggregate collapses twelve different stories into one.

A

Twelve months, one baseline

The same 2,145 held-out rows, split by the month they came from · Skill against the price-only baseline

−0.70−0.60−0.40−0.200+0.15claude-opus-5, w01: below −0.70claude-opus-5, w11: below −0.70oursgpt-5.6-soldeepseek-v4-proclaude-opus-5w00w01w02w03w04w05w06w07w08w09w10w11
w01 & w11 · every outside model losesBelow −0.70 · clipped0 = baseline · higher is betterDecimal skill · 0.01 = 1% of baseline error removed
B

Worst to best partition

Each model’s range across the twelve partitions · The narrow tick marks the mean

ModelPartition skill rangeWorstMeanBest
ours−0.0470.0040.080
gpt-5.6-sol−0.458−0.1100.024
deepseek-v4-pro−0.607−0.1160.069
claude-opus-5< −0.700−0.3000.150
View monthly values
Monthly skill · < means below the display limit
Partitionoursgpt-5.6-soldeepseek-v4-proclaude-opus-5
w00−0.033−0.062−0.140−0.537
w01+0.007−0.458−0.607< −0.700
w02−0.042+0.025+0.014−0.291
w03−0.024+0.018−0.035+0.150
w04−0.029−0.087−0.052−0.406
w05+0.078−0.094−0.100−0.107
w06−0.009−0.005+0.058−0.434
w07+0.046−0.057−0.023+0.042
w08−0.047−0.142−0.031−0.307
w09−0.012−0.030−0.023−0.221
w10+0.080−0.102+0.069−0.084
w11+0.032−0.324−0.521< −0.700
Figure 3 The floor is the result, and the floor has a date. The same held-out rows as Figure 2, split into the twelve months they came from and scored separately, so the horizontal axis is time. Panel B collapses each line into the band it moves in; it is the easier read, but it is the one that loses the shading — two specific months, not two unlucky models.

Two months do nearly all the damage, and they do it to everyone at once. In w01 and again in w11 the three outside models drop together — to −0.46, −0.61 and past the −0.70 clip in the first, −0.32, −0.52 and past the clip again in the second. Three failures that synchronised are not three failures. They are a property of those two months, and the habit being punished is the one Panel B of Figure 2 shows for gpt-5.6-sol: each of these three commits two to three times as far from the baseline as we do. In an ordinary month that costs them a little. In these two it was ruinous. Ours is positive in both, +0.006 and +0.032.

Peaks are cheap. Every model here has a good month, and the largest single win in the figure belongs to the model with much the worst record overall. What none of them has is a floor. Ours moves inside a band roughly a seventh as wide as claude-opus-5's, and on its worst month it gives back 4.7% of the baseline's squared error rather than seventy. Against arriving facts there is no second copy of last March: a month that went badly cannot be rerun. A model that is occasionally brilliant and occasionally catastrophic cannot be left running. Being slightly right most of the time, and barely wrong the rest, is the property that lets a loop keep turning when the outside moves.

The schedule itself is a second finding, not a footnote. At each round we froze a copy and scored it forward, so that training and cadence could be told apart. Training works, and it works immediately: before any of it the base model is 4.14% worse than the baseline, and one window — 461 recorded judgments — takes it to +1.70%. The eleven windows after the first add 4,378 more rows, nine and a half times the data, and move the frozen score only from +1.70% to +1.94%. Per thousand rows the first window is worth roughly fifty times what follows it. Rank correlation and bucket accuracy keep rising across the schedule, so the model is learning something the whole way. It is not converting into squared-error advantage. What is supported here is that the first pass through arriving facts works. That a monthly cadence is what pays is not.

The copy frozen after the first window is scored on eleven later windows, running up to ten months past its freeze point, and shows no decay at all. We read that as a stretch in which the relationship between headlines and repricing held still. We have no independent measure of that, so it is a reading, not a finding. Uncontrolled change is not a promise that every stretch will move. You cannot tell in advance which one will, and the model is the last thing that will tell you.

What repetition does not cost is the thing we were watching for. Plain fine-tuning on a stream of short targets shaves a little off the reasoning on every pass until the model answers in a sentence. Across all twelve rounds the answers stay between 416 and 562 tokens, above the 236-token median of the demonstrations, and after the first window the shortest answer in any batch never falls below 278 tokens. Off-task probes at every checkpoint stay at or above the base model's length, and the prose is normal and correct. The cost is real but narrow: the answer format bleeds into unrelated prompts. Eight to eleven of the twelve off-task answers come back wrapped in the reasoning tags the task uses. The loop can turn again. It does not turn cleanly.

How a verdict becomes a lesson

A wire item lands: a scheduled summit has been postponed. The model reads it and marks the contract on a signed agreement sharply down. It has read “postponed” as “called off”. Three days later the market settles against it: the summit moved by two weeks, and the agreement was signed.

The model that made that call did not announce that it had. Nothing in its own output does, which is why the correction has to come from outside it. Figure 4 is that fact as a picture, not a reading from these twelve months.

TRAINED ON THIS STRETCH OF THE WORLD USED IN THIS ONE the training data ends how people actually behave what the model still assumes error the model's stated confidence flat, the whole way across
Figure 4 Stale by construction. Schematic, not a reading from this run. When the outside moves, the relationship the model learned keeps being applied. Nothing in its own output announces this, which is why the correction has to come from outside it.

The verdict by itself teaches almost nothing. It is one number attached to a noisy problem, and a model trained to reproduce that number learns to copy digits. What we keep is not the number. It is the trajectory: what was visible at the moment of the call, the call itself, what followed, and eventually the settlement — each piece stamped not with when it happened but with when it became knowable, and separated from the call by a cutoff.

ONE COMPLETE TRAJECTORY · TWO TRAINING VIEWS

Record the whole trajectory. Keep the student at the cutoff.

The record includes the evidence, the prediction, every subsequent observation, and settlement. The student’s information cutoff is 14:07.

  1. 13:40

    A wire item

  2. 13:52

    A second item

  3. 14:07

    The model commits

    62% → 28%

    The model lowers its predicted probability.

  4. After the cutoff · as_of = 14:0715:10

    The crowd starts to move

  5. Day 1

    Two related contracts follow

  6. Day 2

    Nothing further

  7. Day 3

    The contract settles

    The judgment was wrong.

Student input: information visible by 14:07.
The prediction is recorded separately as output.

Teacher only: the full continuation, including settlement.
Never added to the student’s original input.

Later, during training

Replay the same prediction with two different views of the record.

Student input

Reconstruct what was knowable at 14:07.

  • News available by 14:07
  • Market state visible at 14:07

No later news, price moves or final outcome.

Teacher context

Use hindsight to guide learning.

  • The same information available at 14:07
  • Crowd movement and related-contract changes
  • Observations with no further change
  • The final settlement

Later evidence can teach the student.
It cannot be added to the student’s original input.

Every record carries three timestamps

event_time
When the event happened.
available_at
When the information first became available.
ingested_at
When our system recorded it.

Evidence arriving after the cutoff is excluded from that prediction’s input.

Figure 5 One complete trajectory, two training views. We record what was knowable at the prediction cutoff, the prediction itself, crowd movement, related-contract changes, observations with no further change, and final settlement. During training, the teacher can use later evidence; the student’s input stays limited to information available at the cutoff.

Let one part of what followed leak back across that cutoff and the outer loop quietly becomes an inner loop: the model is answering a problem we wrote, and nothing in the metrics will say so.

Recording trajectories only pays if you can train on them again without wrecking what you have. Plain fine-tuning on a stream of short targets was the reason a loop that updates forever was not obviously possible. The completion lengths above, holding flat across twelve rounds, are what it looks like when that does not happen.

Reward-based

needs a graderthe world does not supply

Plain fine-tuning, repeated

overwrites what the modelalready had

A conditioned copy of itself

needs neither, and survivesbeing applied again

Figure 6 Why the loop had no engine. Sequential-task experiments show a conditioned copy of itself, the third option here, accumulating skills without regressing on earlier ones (Shenfeld et al., code). That establishes repeated updating need not be self-destructive — not that it makes forecasts better, which is the question this substrate exists to answer.

What replaces it is the model teaching itself. One set of weights plays both parts: a teacher that has been shown how things turned out, and a student that has not.

The same tokens, with and without hindsight

Same model, two context views

  • Student · without later evidence
  • Teacher · with later evidence

Token probability · schematic

a: student 36%, teacher 55%. new: student 46%, teacher 35%. date: student 25%, teacher 60%. is: student 38%, teacher 34%. …: student 24%, teacher 42%.0%40%80%Student: 36%Teacher: 55%aStudent: 46%Teacher: 35%newStudent: 25%Teacher: 60%dateStudent: 38%Teacher: 34%isStudent: 24%Teacher: 42%Hindsight changesthe probabilityTokens sampled by the student

Differences in the next-token distributions guide the student’s update.

Figure 7 What we feed it. The trajectory above goes into the teacher's context and nothing else changes hands. The conditioning can be a bare outcome — on short targets, plain fine-tuning collapses a model's reasoning while a conditioned copy of itself does not (method: Shenfeld et al.).

Nothing larger writes the demonstrations. They come from the student's own weights, shown the continuation and asked to account for it; no stronger model and no human supplies the lesson. Supervision then scales with how much of reality was recorded, not with how much anyone understood.

TrajOps, which we intend to open-source, is the runtime that holds this: it decides what a run was allowed to see, keeps long training resumable, binds checkpoints to the data they consumed, and replays an evaluation against a pinned version months later.

Limits

  1. The cadence finding is above: nearly all of the gain arrives in the first window. This run supports training on arriving facts, not updating monthly.
  2. A fast venue, where the next fact lands within the hour, is a forgiving stand-in for domains where the answer takes years. This tests whether the loop runs, not whether it runs at thirty samples.
  3. The verdicts come from a source we do not own or control.
  4. We model the aggregate belief of a crowd of millions, not millions of individual beliefs.
  5. Each contract is scored on its own. The knock-on case in Figure 1 — a contract that moves only because a linked one did — is the harder problem and is not tested here.
  6. The margin is small and earned on a narrow band. Our model declines to move on most rows and never predicts the largest moves at all, so it would not survive a target with heavier tails than this one.

What this is for

Anyone building a system that predicts collective human behaviour hits the same three walls: what was knowable at the time, when the outcome became knowable, and whether the update actually helped. Those walls are not a research topic, they are plumbing, and they are missing. We would rather build them once, in the open, than watch every team rebuild them badly.

Markets are where we start, because they answer quickly and in public. They are not where this ends. The loop does not care whether the population is human — only that it leaves a timestamped trace and eventually reveals an answer. How millions of people change their minds is one place those traces already exist. It now has an outer loop that turns.