TESSERA v2 reports the largest controlled scaling study for pixel-wise Earth-observation foundation models so far: 395 runs on 1,024 GH200 superchips, evaluated across 15 downstream tasks within a fixed Barlow Twins family. Pretraining loss correlated weakly with downstream performance, with absolute Pearson correlation below 0.2. The proposed compute-allocation rule scales the encoder and data together while keeping the projector fixed. The team then distilled larger models into compact students, including the 21-million-parameter TESSERA v2-1B-M. Its Matryoshka embeddings retain 92% of full 128-dimensional performance at 16 dimensions, according to the abstract.
No heat snapshots are available in the last 24 hours.