Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

BackgroundMellow: A Multi-Modal Framework for Narrative-Driven Cinematic Soundscape Generation

First seen · 7/13/2026, 06:28 PMLatest activity · 7/13/2026, 06:28 PM

BackgroundMellow presents a ground-truth-free framework for generating cohesive cinematic soundscapes from long-form narratives. A master-specialist agent architecture decomposes stories into layered audio cues, assigns categories to suitable specialist models, and combines the outputs through automated mixing. The pipeline uses the Tango2 latent diffusion model for environmental audio and a Cinematic BGM Retriever mined from professional soundtracks. An NLP-based module predicts cue start time, duration, and relative loudness from the narrative timeline. The authors evaluate temporal synchronization, coverage, and spectral richness using nearest-neighbor retrieval against a curated YouTube cinematic-trailer dataset.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/13, 06:28 PMnot independentRepresentative
    BackgroundMellow: A Multi-Modal Framework for Narrative-Driven Cinematic Soundscape Generation