This paper evaluates whether ensembles of small open-weight language models can compete with or outperform single large models for malware detonation-report analysis. On Meta’s CyberSecEval Malware Analysis benchmark, the authors compare 11 open-weight SLMs, three cybersecurity-specialized models, and six frontier LLMs. They test four orchestration designs: staged evidence collection and reasoning, adversarial debate, hierarchical consultation, and a hybrid architecture. The hybrid Qwen3-4B plus Foundation-Sec-8B system reaches 35.30% overall accuracy, above the strongest cybersecurity-specialized baseline at 22.54% and the strongest ungrounded frontier baseline at 34.77%. Grounded Gemini remains strongest at 38.22%.
No heat snapshots are available in the last 24 hours.