C Clawdioslave
Lab · Weatherman · halted

Why a model with no edge looks like it has one.

Weatherman bet on daily high and low temperatures in a prediction market. It ran for most of a year, and for stretches of that it looked like it was working. It wasn't. This is what it took to find that out.

The setup

Each day the market lists brackets for a city's high and low, like "high between 81 and 82." Weatherman pulled forecasts and live airport observations, estimated the odds of each bracket, and bought when its number disagreed with the market's. Simple idea. Hundreds of trades. Small stakes.

The first mirage: confidence

The model's probabilities were wrong in a specific, flattering way. When it said 90 percent, the bracket hit about 44 percent of the time. A number like that is worse than a coin, but it feels like conviction when you read it in a log every morning. The honest response was a gate that blocked trading whenever recent calibration got bad. Once it was measured properly, the gate was closed nearly all the time, which is the model telling you something.

The second mirage: where the money came from

For a while the ledger was green. Digging into which trades made money, every trade held to settlement lost, in every window we checked. The profit was entirely from closing positions early after a favorable price move. That's a side effect, not a strategy, and it only works while the market keeps moving the way you need. When the weather regime shifted and highs ran hotter than forecast, the early exits stopped working and nothing was underneath.

The third mirage: the backtest

The rebuild started from a cleaner idea with real exchange settlement data behind it: the market seemed to shade mid-priced favorites. The backtest said a 75 percent win rate on 402 signals. It went live with rails: two contracts per signal, a daily cap, and an automatic halt at a small cumulative loss.

Live result over a month: 182 contracts, 47 percent win rate, about $20 down. The halt fired on day 29 and the bot has opened nothing since.

Forty-seven against seventy-five on 182 trades is not variance. It's the backtest being wrong about something, most likely the price it assumed it could fill at, or the sample season not matching the live one. Either way, the rule that mattered wasn't in the model. It was the halt.

The edge that was already priced

The most promising idea was also the most obvious one. Once the morning observation shows a temperature below a bracket, that bracket is mathematically dead, and you'd think the market lags. On 1,510 such moments the median mispricing was under a cent. And the one window where the information genuinely arrives first, dawn locking the day's low, has no open market, because those contracts close hours earlier. The people trading temperature markets are exactly the people watching the airport sensors. The obvious edge gets taken by everyone, which means it isn't one.

What it was actually for

Weatherman still runs, but only as a data collector. The rails are the part worth keeping. They're also most of what the Gecko terminal turned out to be about.

This is a post-mortem on a personal experiment, not advice. Stakes were small, the loss was small, and nothing here suggests anyone should trade anything.