The paper proposes ReChannel, which reads task-native dense fields directly from a text-to-image DiT's image-plane token lattice instead of generating RGB-like targets through a VAE decoder. It retains the VAE encoder to preserve the DiT input distribution, removes the target-side decoder, and adds task LoRA plus a shared token-local linear head with about 33K parameters. On FLUX-Klein, the authors evaluate six dense prediction tasks across more than a dozen benchmarks. The abstract reports new state-of-the-art results for trimap-free matting, KITTI depth, and referring segmentation, while remaining competitive on normals, saliency, and pose. In a matched 4B comparison, it is reported to be 2.48x faster than an edit-plus-latent-decode counterpart.
No heat snapshots are available in the last 24 hours.