What’s new in TEE: global coverage, Deep MLP, and spatial splits
TEE, the Tessera Embeddings Explorer, lets you browse Tessera embeddings for any 5 km × 5 km patch of Earth, label habitats, and evaluate classifiers on your own ground truth — all without writing code. I’ve added three new features in the last month, before I dive into an intense teaching term (so, a freeze on features).
Embeddings from 2017–2025
TEE’s embeddings now come from v1.1-dclimate, a wall-to-wall Tessera dataset covering the entire land surface of the planet for every year from 2017 to 2025. Previously, coverage was only for 2024 for most of the world. In practice this means that if you see “No GeoTessera tiles found”, it’s because you’re over the sea or outside 2017–2025, not because your region of interest fell into a hole in the data. (A handful of areas are still missing in individual years for lack of satellite observations — try a neighbouring year if that happens.)
Deep MLP: a heavier classifier option
Validation now offers Deep MLP (BatchNorm) alongside k-NN, Random Forest, XGBoost, MLP, Spatial MLP, and U-Net. It’s a PyTorch network with batch normalization, dropout, and the AdamW optimizer, which is the same architecture the original Tessera paper used for its own downstream evaluations. It generally beats plain MLP given enough training data. It needs PyTorch on the compute server, the same as U-Net.
Dealing with spatial leakage?
The biggest additions are the following options that prevent a classifier “cheating” by exploiting spatial correlation rather than actually learning the habitat:
- Group by field keeps every pixel from one shapefile polygon entirely on one side of the train/test split. Without it, pixels from the same field can land on both sides, letting a model partly recognize the field instead of the habitat. For the ~2M-pixel Austrian dataset this reduced macro-F1 by 0.03.
- Spatial k-fold does the geographic equivalent for k-fold cross-validation: each point is assigned to one of k geographic blocks, and no block’s points cross between folds. Plain k-fold shuffles points randomly with no notion of location.
- Area-stratified split — 30% of each class’s field area goes to training, independently per class, which we used in the Tessera v1 paper. Comparing the options side by side on the same dataset gave macro-F1 of 0.79 (default), 0.68 (group by field), and 0.61 (area-stratified), a reminder that split methodology alone can move a score as much as the model does.
All three are off by default, and are documented in the User Guide.
Smaller things worth knowing about
- Spatial MLP’s training-point and per-patch limits are now visible, settable fields instead of hidden constants.
- Point ground truth (not just polygons) now works throughout Validation.
Try it at tee.cl.cam.ac.uk, or self-host from github.com/ucam-eo/TEE.