Accuracy figures on this page are withdrawn pending recompute
The per-crop correlations and the overall headline below are in-sample figures: the model was scored partly against county-year cells it was fitted on. They are presented here as an out-of-sample breakdown, and that is wrong. The overall figure also pools eight crops whose yields span roughly 2,000 to 28,000 kg/ha, so most of the correlation is the spread between crops rather than skill within one — predicting each crop's average would score about the same.
A proper leave-one-year-out, per-crop evaluation is being recomputed, together with a headline measured against a per-crop-mean baseline. Until that lands, please do not quote any number on this page. The map, the county detail and the scatter are unaffected; only the accuracy claims are.
Flagged 15 September 2026 · notice added 21 September 2026.
County yield: predicted vs INSSE
Per-county crop yield (kg/ha) estimated from TESSERA 128-dim
embeddings (centroid-sampled per field, area-weighted to
crop-county-year), then a LightGBM regressor trained on the
log target with crop + county + year fixed effects. Predictions are
compared against INSSE TEMPO matrix AGR110A (average
yield per hectare). Validated under leave-one-year-out cross
validation: model lift +0.071 R² over the
crop+county+year baseline, well above the +0.05 spec gate.
Methodology note. The model is supervised at county level only — per-field values are downscaled from the county-trained model with aggregate-consistency calibration (clip 0.7–1.4×). They are not field-level ground truth.
Agreement scatter — selected crop, all years
Each dot = one (county, year). Diagonal = perfect agreement.Model performance
LOYO R² on log(yield) per the lift gate experiment. This card is the leave-one-year-out comparison and is not among the withdrawn figures.Why TESSERA? The variant comparison evaluated AEF (64-dim), TESSERA (128-dim), and dual (AEF+TESSERA) on the same baseline ladder. AEF alone reached +0.027 lift over B3 (failed the gate), TESSERA reached +0.071 (passed), dual added zero marginal gain. TESSERA's self-supervised representation appears to capture phenology and canopy state more densely than AEF's spectral-temporal encoding for yield-relevant variation.
How to read these numbers
Per-crop Pearson r is meant to measure how well the model ranks counties for that crop — given the actual crop, does the predicted yield track INSSE county-to-county? The per-crop values shown here are in-sample and are withdrawn pending recompute (see the notice at the top), so the ordering they suggest — oilseeds and cereals strong, potatoes and wheat weak — should be treated as unverified rather than as a finding. The stated reason for the weaker crops, that they depend on management factors such as irrigation, cultivar and sowing date that the embedding does not directly see, remains a plausible expectation but is not evidence from these numbers.
MAPE is the mean absolute percentage error across all (county, year) cells for the crop. Low single-digit MAPE means predictions land within a few percent of the INSSE value; double-digit MAPE means the model gets the direction right but misses the magnitude.
The overall headline is withdrawn. It was reported as an out-of-sample correlation across held-out cells, but it pools all eight crops, so it largely reflects the scale difference between them rather than the model's skill within a crop. A within-crop measure against a per-crop-mean baseline is being computed to replace it.