Modus explores decoder-only any-to-any modeling, using one network to predict any modality from arbitrary combinations of other modalities. It treats every modality symmetrically and avoids modality-specific heads, losses, and task pipelines, potentially allowing strong pretrained decoder-only models to serve as the prior. The authors describe applications including generation chained through intermediate modalities and cross-modal self-verification, where outputs are scored through another generated modality. According to the abstract, one Modus model is competitive with specialist and multitask baselines across several benchmarks. The project materials are released publicly by EPFL.
No heat snapshots are available in the last 24 hours.