The paper introduces the Authorship-Rewriting Benchmark (ARB), a matched dataset built from 1,800 human texts drawn from XSum, WritingPrompts, and OpenWebText. Each source produces four variants: human-written text, direct LLM generation, human-to-LLM rewriting, and same-generator rewriting of LLM text. At a strict 1% false-positive rate, FastDetectGPT and Binoculars-falcon-7b detected 91.2% and 93.5% of direct LLM text, but only 30.8% and 15.1% of human text rewritten by an LLM. Detection of LLM-originated text remained substantially stronger after rewriting.
No heat snapshots are available in the last 24 hours.