HOMIE is a unified framework for human-object centric video personalization under both inter-subject and intra-subject reference settings. It uses global multimodal guidance inside self-attention to align semantic features extracted by a multimodal large language model with VAE tokens, while modality-reference embeddings distinguish MLLM and VAE tokens and associate intra-subject reference image tokens. The authors report state-of-the-art results across several HOCVP tasks, targeting improved subject fidelity and human-object interaction, including abstract objects such as logos.
No heat snapshots are available in the last 24 hours.