The paper introduces CodeTracer, a forensic framework for tracing malicious code completions back to the fine-tuning data that implanted a backdoor. Under post-deployment constraints, it uses only the fine-tuning corpus and a reported miscompletion event. CodeTracer extracts a structured behavioral fingerprint from the unsafe output, filters for semantically relevant code samples, and uses LLM-based reasoning to identify likely responsible data. The authors report evaluations covering three vulnerability cases, ten backdoor attacks, and sixteen competitive baselines, with high forensic accuracy, low false-identification rates, and robustness against adaptive attacks.
No heat snapshots are available in the last 24 hours.