This paper presents a harness-centered system for LLM-driven GPU kernel optimization on NVIDIA Blackwell B200 GPUs, developed for the MLSys 2026 FlashInfer AI Kernel Generation Contest. The system separates an evaluation harness from a profile-backed optimization controller. The harness checks compilation, correctness, official-aligned timing, and artifact archival, while the controller uses profiler and workload evidence to bound candidate generation and promotion. Across five operator definitions, retained artifacts achieved mean-latency speedups of 1.62x, 18.05x, 29.68x, 1.12x, and 13.70x over supplied FlashInfer baselines. Agent-Assisted kernels outperformed Full-Agent artifacts, emphasizing the continuing importance of expert constraints, references, profiling procedures, and workload context.
No heat snapshots are available in the last 24 hours.