This paper studies culturally loaded machine translation through a Chinese-Japanese dataset built from Dream of the Red Chamber. The dataset contains 500 segments spanning diverse cultural categories. Using a comprehensive evaluation protocol, the authors identify three systematic problems: frontier LLMs still show notable performance gaps on culturally embedded content; human judgments can differ substantially because evaluators have different cultural backgrounds; and widely used automatic metrics do not reliably measure translation quality in this setting. The work frames cultural translation as an evaluation problem as well as a modeling problem.
No heat snapshots are available in the last 24 hours.