This paper reports official Task 1 results and development diagnostics for LongEval-Sci 2026, a benchmark for scientific retrieval under collection change. It compares PyTerrier BM25 and Qwen3 dense baselines with full-text BM25, temporal and citation features, RM3, cross-encoder reranking, and reciprocal rank fusion. Temporalized full-text runs achieved the best ARP across all three official snapshots, with nDCG@10 values of 0.285, 0.267, and 0.180, while reducing snapshot-3 relative change from 0.481 to 0.368 against the BM25 pivot. Full-text BM25 was strongest on internal snapshot-1 diagnostics, while RRF delivered the best Recall@1000.
No heat snapshots are available in the last 24 hours.