Bellingcat describes a machine-learning workflow for identifying civilian-harm reports in Telegram data related to the war in Ukraine. According to the article title and supplied metadata, XGBoost outperformed the large language models tested for this classification task. The report is relevant because it connects model selection with real investigative monitoring needs, where structured features, reproducibility, cost, and review capacity may matter as much as general language ability. The supplied summary does not include the benchmark design, model names, dataset size, or evaluation metrics.
No heat snapshots are available in the last 24 hours.