Evidence
We try to catch ourselves lying.
Evidence, not adjectives. Every number below renders from metrics.json and carries its ratification status — including the arms that lose.
Current gate table — 2.5 m plane
| Model | KID-HP ↓ | DISTS ↓ | LPIPS ↓ | PSD ↓ | P16 | Measured on | Status |
|---|---|---|---|---|---|---|---|
| Bicubic ×4, radiometry-normalized | 18 297 | 0.342 | 0.720 | 1.594 | 0.90 | bicubic | Control |
| SEAN v5 | 3 262 | 0.145 | 0.225 | 0.100 | 1.00 | outputs/sean/v5/sean_best.pt | Ratified |
| SEAN v6 without the projected discriminator | 3 503 | 0.142 | 0.224 | 0.097 | 1.00 | outputs/sean/v6_noprojd/sean_best.pt | measured |
| SEAN v6 | 1 959 | 0.134 | 0.225 | 0.095 | 1.00 | outputs/sean/v6/sean_best.pt | Pending sign-off |
| Demo default — v6 scale run @58k | 1 622 | 0.130 | 0.220 | 0.089 | 1.00 | outputs/sean/v6_scale120k_projd/sean_058000.pt | Ratified · demo use |
| PiSA-SR zero-shot | 7 777 | 0.238 | 0.426 | 0.337 | 0.98 | pisasr_zs | External |
| OSEDiff zero-shot | 9 302 | 0.243 | 0.465 | 0.425 | 0.98 | osediff_zs | External |
| SEN2SRLite | 19 117 | 0.371 | 0.736 | 1.433 | 1.10 | [object Object] | External |
| LDSR-S2 | 19 175 | 0.343 | 0.586 | 1.033 | 1.02 | [object Object] | External |
| DiffFuSR | 22 757 | 0.296 | 0.447 | 0.316 | 0.99 | [object Object] | External |
Real-input arms: LPIPS and PSD are content-confounded, because the reference photograph was taken on another date. KID-HP and DISTS carry the decision. The 0.5 m research track uses a different fixture and scale and is deliberately not shown beside these numbers.
What each column means
- KID-HP ↓
- Kernel Inception Distance on high-pass-filtered patches. Does our fine texture look statistically drawn from real aerial photography? The primary realism gate.
- DISTS ↓
- Perceptual structure-and-texture distance to the reference photograph. A second opinion, more tolerant of texture resampling than LPIPS.
- LPIPS ↓
- Perceptual distance to a reference photograph captured on a different date — not ground truth, and not a target we aim at. Recorded, but demoted from decisions for exactly that reason.
- PSD ↓
- Radial power-spectral-density distance. Is detail present in the right amount at every scale, rather than piled into one frequency band?
- P16
- Blockiness score at a 16-pixel pitch. 1.00 means no periodic block artifact from the generator's tiling.
- 2AFC
- Can a human tell ours from real? 0.50 = coin flip = the goal. The actual success criterion, and the one we are furthest from.
Scope of the evidence
Still measured on the previous model
We changed the model before we re-ran every check on it. These four numbers describe SEAN v5, not the checkpoint the imagery on this site comes from. There is no suite score for the current model on real Sentinel-2 input, so the comparison against public models cannot yet be re-based, and no round-trip run exists for any demo checkpoint.
| Evidence | Value | Measured on | Date |
|---|---|---|---|
| Comparison against public models | v5 4 630.7 vs PiSA-SR 7 777.3 | v5_real_raw | 2026-07-05 |
| opensr correctness triad | hallucination 0.284 / omission 0.506 / improvement 0.210 | v5_real_raw, 24 tiles | 2026-07-05 |
| Founder 2AFC self-test | 0.90 (goal 0.50) | v5_best / a10_product_cal | 2026-07-06, 07-11 |
| Round-trip at the demo checkpoints | never run — v6 @31.1k is the nearest point, at 0.8978 | v6 @31.1k | 2026-07-07 |
The number we would rather not show you
The same model, measured two ways
A conditioning map tells the generator where things are. In training, part of it was computed from the very photograph we were trying to match — so it already knew some of the answer. That gives an optimistic result. Switch to deployment-honest for the same checkpoint scored with conditioning we could actually build from Sentinel-2 and OpenStreetMap alone. It is the valid practical comparison, and it is the one we lead with.
| Stratum | KID-HP ↓ deployment-honest | KID-HP ↓ oracle | |
|---|---|---|---|
| All strata pooled | 2 896.8 | 1 622.2 | |
| Residential | 3 331.9 | 1 836.0 | |
| Industrial / port | 3 837.0 | 1 918.9 | |
| Agricultural rows | 4 144.3 | 2 348.9 | |
| Forest | 1 877.2 | 1 819.5 | |
| Water / coast | 2 867.9 | 1 738.6 |
The model was never trained on deploy-S maps, so the deploy column is off-manifold and a leak-free retrain would likely land between the two columns. Forest is nearly leak-independent; urban and agricultural strata ride on the oracle vegetation extent.
Leaked: max(OSM perennial landuse, ExG from hr_2.5) — HR-derived; annual cropland gets NO OSM burn, ExG alone decides (build_sv2.py:234-241). Not leaked: road, road_service, building, rail, water and wall come from OSM vectors; the NDWI ExG-suppressor is LR-available. Extent: ~68% of vegetation-labeled pixels on average are image-derived; ~97-99% in dense urban Koper (reports/demo/final/CLAIMS_AND_LIMITS.md). Next: Leak-free vegetation conditioning retrain is scoped and not started (user-gated) — reports/demo/final/CLAIMS_AND_LIMITS.md
Checkpoint selection
Why this checkpoint, and what it cost
Stage-A realism per stratum across the demo training run. The ruling is not '58k is best' — pooled is flat from 58k while industrial_port regresses after it.
| Stratum | KID-HP ↓ | |
|---|---|---|
| All strata pooled | 1 622.2 | |
| Residential | 1 836.0 | |
| Industrial / port | 1 918.9 | |
| Agricultural rows | 2 348.9 | |
| Forest | 1 819.5 | |
| Water / coast | 1 738.6 |
58k is the demo default; 100k is appendix-only (user, 2026-07-28). Rationale recorded: pooled plateau by 58k plus a substantial industrial regression at 100k for modest gains elsewhere. Final visual ratification: PENDING HUMAN REVIEW.
forest = 19 crops; small-n noisy.
One variable
What the discriminator is worth
Two runs, one difference. With the projected discriminator, 1 959.2; without it, 3 502.6 — against 3 262.0 for the previous generation. Removing it costs essentially the whole improvement. Both arms are measured and pending sign-off: the permanent champion swap is unratified, and the recorded GO was a delegated model read, explicitly not human-verified.
Texture realism per version
The best number on this chart is still pending sign-off; it is not treated as shipped.
Against public models (identical inputs)
≥1.7× on the strongest public model. The Sentinel-2-native public models sit at the bicubic floor.
Can you tell which is real?
Twenty pairs. One patch is real aerial photography, one is generated from a Sentinel-2 pass. Chance is 0.50 — that is the goal, and we are not there yet: the founder's logged sessions all scored 0.90.
| Session | Model | Answered | Accuracy · goal 0.50 | Calibration FPR | Graded |
|---|---|---|---|---|---|
| 1 | v5_best | 20 | 0.90 | 0.00 | 2026-07-06 09:20 |
| 2 | a10_product_cal | 15 | 0.90 | 0.40 | 2026-07-11 09:59 |
| 3 | a10_product_cal | 15 | 0.90 | 0.40 | 2026-07-11 11:02 |
Sessions 2 and 3 are one session, graded twice and logged twice. Shown as logged rather than deduplicated. The calibration false-positive rate rose from 0.00 to 0.40 between sittings, which is a caveat on the accuracy beside it. All three were run on the previous model — see “Still measured on the previous model” above.
1 / 20
Session complete
The goal is 0.50 — a coin flip. Anything well above it means the generated patches are still telling on themselves, and that is the number we are trying to move.
Your running accuracy — dashed line = 0.50 goal
no answers yet
Your session is kept in this browser only. Nothing is sent anywhere.
Method & data shipping today
- Corpus: 730 co-registered Sentinel-2 ↔ aerial pairs over 7 granules, ≈1,196 km², with 10-frame temporal stacks per site.
- Alignment: every revisit is registered to a common grid; misregistered frames are dropped, not averaged in.
- Structure map: S_struct, a 6-class blueprint — roads, buildings, rail, water, vegetation, bare — carried with a confidence channel.
- Generation: SEAN generator + OASIS discriminator GAN, structure-conditioned, 10 m → 2.5 m (×4).
- Verification: round-trip to 10 m plus texture statistics against real-aerial crop banks, on every output.
- Publication: the gate table, the caveats and the failures ship together, from one file.
Target architecture in development
- Multi-temporal conditioning: all ten Sentinel-2 revisits enter the generator instead of one date-matched frame.
- Two control branches — structure and the date-matched low-resolution pass — conditioning a diffusion decoder.
- Stage B, a 2.5 m → 0.5 m texture octave research-only · NSCLv1 weights
- Height estimation from the generated plane, published only once the round-trip check covers elevation.
- Live inference for requested areas, replacing the current send-it-to-us route.
ControlNet conditioning was tested and concluded negative in 2026-06 experiments; this is a target, not a shipped result, and it is marked that way on purpose.
Against ourselves
The ones we lost
Claims we withdrew, tests we failed, and numbers that turned out to be wrong. Self-published, unprompted, and kept above the report shelf while that shelf is empty.
- 01The LLM judge is UNVALIDATED (repeat-agreement failed both prompt versions) — its numbers are recorded facts, never decision inputs.
- 02Two historical hf-fraction definitions exist; only the pinned def-A values with a stated tile set are citable.
- 03The honest bicubic floor is LPIPS ≈ 0.720, not the 0.841 recorded before radiometric normalization.
- 04A 2026-06-21 claim that real-input output HF was identical to synthetic input was withdrawn — the deficit is ~15% and was masked by rounding.
- 05Indistinguishability is a goal, not an achievement: the founder still tells our output from real aerial 90% of the time.
- 06LLM-judge repeat agreement: 0 (threshold 0.8) — UNVALIDATED — recorded facts, not decision inputs.
- 07ControlNet conditioning was concluded NEGATIVE at both latent and pixel sites (2026-06-19); we pivoted to a SEAN+OASIS GAN.
- 08High-pass spectral fine-tuning scored best on the realism axis and was still closed as a Stage-A candidate (2026-07-11).
- 09The 2.5 m HF-adversarial refiner regressed at every checkpoint in both arms (2026-07-06) — 3 262.0 → 4 608.3 and 2 878.1 → 4 541.3.
- 10The A03 mush-inversion probe fired its own pre-registered kill rule and the residual critic was demoted to hygiene use (2026-07-10).
- 11Both S-generation fold-0 cross-validation arms hit the pre-registered mid-run kill at 10.5k steps; the validation gate was never reached (2026-07-21).
- 12Our conditioning map was partly computed from the photograph we were trying to match — audited 2026-07-27, and the honest re-score is above, not hidden.
Reports
The baseline table, the metrology note and the validity annex exist internally. None is published until the founder's release selection is made, and this shelf stays visibly empty until then. The numbers those reports pin are the ones already in the table above.