ArbDeskv4
Polymarket daily-temperature markets · forecast vs. book
Bankroll
Open exposure
Day P&L (net)
UTC --:--:--
Next model cycle
Vol 24h

Predictive

The same question in both directions. Forward: what the desk expects, which bucket that lands in, and what the market is charging for it. Backward: when it expected that before, was it right — and separately, did being right pay. Those are different questions with different answers, because entry price, fees and fill size all sit between them.

What the desk expects

One row per city-day still open. The bucket the model puts most weight on, what the market charges for that bucket, and the largest tradeable edge anywhere on the ladder. A row with no probability has a forecast but no priced market yet.

Loading…

How the forecast converged

Depth is how many days before resolution, height is temperature, width is the resolution day. Each ribbon is one model walking from what it said a week out to what it said that morning; the green bar is what actually happened. A good model's ribbons narrow toward that bar. A biased one runs parallel to it and never arrives — and those two look identical in a table of mean error. The shaded planes are the real bucket boundaries, because a 0.6 °C miss across a line loses and a 0.9 °C miss inside one wins.

Loading…

Actual against predicted

the only question that decides whether the rest of this page is worth reading

Every settled day, one dot: what was forecast a day out against what the day actually did. On the diagonal is a perfect call; above it the day came in hotter than said, below it cooler. Green is within 1 °C, which is roughly one bucket — the resolution the market actually pays at, so a dot being green matters more than it being close. Bias and error are read separately: a forecast that is 1.5 °C hot every day is a correction you can apply, and one that is 1.5 °C off in random directions is not. The numbers below split them.

Loading…

How fast the forecast decays

The same error, split by how far ahead the call was made. A model that is sharp tomorrow and useless on Friday averages out to “fine” — this is the only place that shows. The practical use is picking the lead day at which to stop trusting it: where the line crosses one bucket wide, the forecast has stopped resolving which bucket wins.

Loading…

Hit rate, per city, per lead

Bias is kept separate from error on purpose: a model 1.5 °C hot every single day is fixable, one that is 1.5 °C off in random directions is not, and pooling them into “1.5 °C error” throws away which you have. Hit rate is the only accuracy the market pays for — did the day land in the bucket the forecast pointed at.

Loading…

Bankroll

Realised P&L, cumulative, from filled paper trades only — a proposal that never filled cost nothing and proved nothing. The win rate beside it is running, not final.

Loading…

Can it scale?

Claimed edge against what that edge actually returned. If realised tracks claimed, size can go up. If realised is flat whatever was claimed, the edge estimate is noise — and sizing up multiplies noise, not profit.

Loading…