This survey reviews computational humor for multimodal large language models, focusing on memes, cartoons, comics, and both single-image and multi-panel artifacts. It organizes prior work around recognition, interpretation and reasoning, and generation. The authors trace a shift from task-specific fusion systems toward multimodal alignment, evidence-grounded reasoning, and controlled generation. The survey identifies shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns as central obstacles. Humor generation is presented as an emerging downstream frontier rather than a mature capability.
No heat snapshots are available in the last 24 hours.