Analytics
Four questions, in the order a desk actually asks them: is the forecast any good, is the pricing any good, did it make money, and can it take size. Every chart reads the same tables the trading pages do — nothing here is a second calculation of something computed elsewhere. The first two groups need no trades at all; the third and fourth are marked, so an empty desk reads as early rather than broken.
Is the forecast any good?
What the archive can say about itself, with no trade ever placed. If these are weak, nothing below them can be strong — every band probability is built on this.
Calibration — is a 30% actually a 30%?
Every band the desk priced, grouped by what the model claimed, against how often it happened. A model can have a wonderful error and still be systematically overconfident, and that shows up here and nowhere else.
What am I looking at? →Hide the explanation ↑
The question. The desk does not predict temperatures, it prices probabilities: “this band has a 30% chance”. A probability is only worth anything if it is honest — if the things it calls 30% happen about 30% of the time. That is what calibration means, and it is a completely different question from being accurate.
How to read it. Every band the desk has ever priced is put in a bucket by what the model claimed, then we count how many of them actually happened. Take the bucket where the model said 30%. If 30 out of those 100 bands won, the dot sits on the diagonal and the model is honest at 30%. If only 18 won, the dot sits below the line: the model says 30% and means 18%, so it is overconfident and every edge computed from it is overstated. If 42 won, it is underconfident and the desk is leaving money on the table.
Why it is not the same as being accurate. A model can have an excellent temperature error and still be badly calibrated — right about the middle and far too sure of itself about the spread. That failure is invisible in every error metric and shows up here and nowhere else. When it appears, scripts/calibration.py fits a correction and the desk stops overstating its edge; sql/ad4_45 does the same for the width of the distribution.
What to do about it. On the diagonal: nothing, trade the edges as computed. Consistently below: the desk is over-betting and every position should be smaller until the correction lands. Consistently above: it is under-betting. A curve built on fewer than 300 settled outcomes is noise, so it is not drawn at all rather than drawn faintly.
Does claimed edge become money?
Edge is a prediction about profit. This is the only place it is checked against what actually happened — grouped by how much edge was claimed, so a desk that is right about small edges and wrong about large ones can see it.
What am I looking at? →Hide the explanation ↑
What “edge” is. When the market prices a band at 30c and the desk's model says it has a 38% chance, the desk is claiming 8 percentage points of edge: it believes it is buying something worth 38c for 30c. That number is a prediction about profit, and like any prediction it can be wrong.
What this table checks. Every settled band is grouped by how much edge was claimed on it, and then we ask what those bands actually returned. Claimed is what the desk said it was getting. Realised is what it got. Buy a hundred bands at 30c claiming 8 pp of edge: if 38 of them win, you paid $30 and collected $38, and realised matches claimed. If only 31 win, you collected $31 — the claim was 8 pp and the reality was 1 pp, and seven points of the edge were imaginary.
Why it is split by size. The interesting failure is not being wrong everywhere — it is being right about small edges and wrong about big ones. A 2 pp claim that realises 2 pp and a 20 pp claim that realises −4 pp average out to something that looks fine, and they mean opposite things: a huge claimed edge is usually the model misunderstanding a market rather than beating it. Split by bucket, that shows immediately.
What to do about it. Realised tracking claimed down the column: the model is real, trade it. Realised falling away in the large-edge rows only: cap the size the desk will take on a big claimed edge, because those are the ones it is wrong about. Realised negative across the board: the edge is not there and the fee and slippage model is eating it — check the fill assumptions on Goals before the model.
AD4's forward model against NWS and persistence
Per LEAD DAY, because a model that is sharp today and useless on Friday looks fine averaged together. Persistence — yesterday's maximum, unchanged — is the bar; NWS is the forecast this model post-processes.
How wrong is the CLOUD forecast?
The desk measures temperature error and has never measured this. If cloud is worth about −1 °C per okta, a cloud forecast two oktas out is a 2 °C error entering the model through a side door — invisible to every temperature-based skill metric.
What actually moves a day
Measured on this desk's own archive, in plain buckets: how much further the afternoon climbed from its morning reading under each condition. This is where the model's coefficients come from, before any fitting.
Persistence — the bar, per city
Yesterday's maximum, unchanged. It is not a strawman: in a stable air mass it is very hard to beat, and a city with a small day-to-day swing is a city where no forecast can add much. Sorted by how beatable each city is.
How much climb is left, by hour
Measured per city and local hour: how much further the day still climbed from here, historically. This is what turns a rate of change into a decision — +0.9 °C/h at 13:00 is ordinary in one city and remarkable in another.
Forecast skill by city (1-day lead)
Mean absolute error against what actually happened — the number sigma is built from, so a city high on this chart is one where every band probability is necessarily vague. This is the one-day view; Predictive has the same measure per lead day, with bias and hit rate beside it, which is where to go when a city looks wrong here.
Is the pricing any good?
The forecast turned into a probability, and the probability against what the market charges. Still no trades required — this is the model and the book disagreeing on paper.
Model against market
Every tradeable YES band: what it costs against what the model thinks it is worth. The diagonal is agreement. Above the line the model is more optimistic than the book — that is a buy, and the vertical distance is the raw edge before costs. Below it, the market is paying more than the model thinks the band is worth. Colour is the city's hotness against its own normal, so a cluster of buys on an unusually hot day is visible as one.
What the book costs you, per city
Exactly one bucket pays, so the YES prices across a city's ladder should sum to about $1. What they actually sum to is the overround — the house's margin, and the first thing subtracted from every edge on that board. Under $1.00 is not a rounding artefact: it is a combination arb, which is s2's entire thesis and needs no forecast at all. Beside it, the median spread you cross to get in, and how much of the ladder the edge engine will actually let you trade.
Monte Carlo — what a position actually does
The engine already computes each band's probability analytically, so simulating the same normal would only reproduce it. What is not obvious is the shape of a position's outcomes: several legs, a fee that varies with price, and an all-or-nothing payoff per band. A basket with positive expected value that loses four times out of five is a different proposition from one that loses one time in five, and the average cannot tell them apart. 20,000 draws of the day's maximum from the model's own centre and sigma — the same distribution the desk is betting on.
v_opportunities to have rows.Did it make money?
The first group where a trade has to have happened. On a desk with every strategy switched off these are empty, and that is early rather than broken - the Strategies page is where that changes.
Realised P&L
Cumulative net profit from closed paper trades, in settlement order. Net means after fees and spread — the desk quotes nothing gross.
Strategy attribution
The desk's own record
Every other chart on this page reads live tables that are overwritten by their own next run. The frozen facts — what was predicted, what the market charged, what actually happened — now have their own page, along with the archive underneath them and every derived layer built out of it.
Open the Data Bank →Can it take size?
An edge you cannot fill is not an edge. Needs book snapshots from P0.3 and traded volume from P0.4.
Liquidity: quoted against traded
Two different facts, plotted rather than merged. Depth is what the current quotes can absorb inside 5¢; volume is what actually changed hands. Up and to the right is a real market. Top-left is quoted but not traded — a fat quote nobody hits, and the opportunity ranking already discounts it. Bottom-right trades in bursts against a thin book, which costs money on the way out rather than the way in.