This paper benchmarks seven models for AWS Terraform generation across 17 scenarios: Claude Opus 4, GPT-5.4, Gemini 2.5 Pro, Qwen2.5-Coder-14B, WizardCoder-33B, CodeLlama-13B, and Magicoder-S-CL-7B. The study combines Checkov and Trivy with a GitLab CI/CD pipeline and evaluates three security prompt levels using pass@5. Its central finding is that syntactic validity and security compliance are largely orthogonal. WizardCoder-33B reaches a 77.8% validation rate but zero Checkov compliance, while Claude Opus 4 reaches 23.1% Checkov and 92.5% Trivy pass rates with detailed security prompting. The authors conclude that automated multi-tool scanning remains necessary regardless of model family or prompt strategy.
No heat snapshots are available in the last 24 hours.