The paper introduces HuGLEN, a stepwise evaluation pipeline that combines an LLM-as-a-judge with a small number of expert ratings to compare candidate language models for network automation. It ranks models using a quality efficiency score (QES), incorporating explanation quality and inference efficiency. The demonstrated task translates outputs from an explainable AI model for optical-network quality-of-transmission (QoT) estimation into operator-friendly explanations. According to the abstract, a medium-sized 12B-parameter model achieves the highest QES, suggesting that the largest model is not necessarily the best operational choice.
No heat snapshots are available in the last 24 hours.