BackgroundMellow presents a ground-truth-free framework for generating cohesive cinematic soundscapes from long-form narratives. A master-specialist agent architecture decomposes stories into layered audio cues, assigns categories to suitable specialist models, and combines the outputs through automated mixing. The pipeline uses the Tango2 latent diffusion model for environmental audio and a Cinematic BGM Retriever mined from professional soundtracks. An NLP-based module predicts cue start time, duration, and relative loudness from the narrative timeline. The authors evaluate temporal synchronization, coverage, and spectral richness using nearest-neighbor retrieval against a curated YouTube cinematic-trailer dataset.
No heat snapshots are available in the last 24 hours.