The paper presents dots.tts.edit, a speech editor adapted from the continuous autoregressive dots.tts model. Instead of ambiguous free-form prompts or explicit timestamps, it uses XML-style structural instructions grounded in transcript spans and boundaries. The system covers text, emotion, pitch and speaking rate, and pause editing. Its bilingual doteBench suite evaluates instruction following, preservation outside edited regions, audio quality, and composed operations. The abstract reports leading overall instruction following and local preservation across five editing categories, with quality comparable to open-source systems; code and model release is still pending.
No heat snapshots are available in the last 24 hours.