Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Zhe, Liang, Tuo, Li, Jing, Lu, Yiren, Zhou, Yunlai, Qiao, Yiran, Ma, Jing, Yin, Yu |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
par: Liang, Tuo, et autres
Publié: (2025)
par: Liang, Tuo, et autres
Publié: (2025)
Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting
par: Lu, Yiren, et autres
Publié: (2025)
par: Lu, Yiren, et autres
Publié: (2025)
CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data
par: Liu, Disheng, et autres
Publié: (2025)
par: Liu, Disheng, et autres
Publié: (2025)
AdvSplat: Adversarial Attacks on Feed-Forward Gaussian Splatting Models
par: Qiao, Yiran, et autres
Publié: (2026)
par: Qiao, Yiran, et autres
Publié: (2026)
DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering
par: Qiao, Yiran, et autres
Publié: (2026)
par: Qiao, Yiran, et autres
Publié: (2026)
BARD-GS: Blur-Aware Reconstruction of Dynamic Scenes via Gaussian Splatting
par: Lu, Yiren, et autres
Publié: (2025)
par: Lu, Yiren, et autres
Publié: (2025)
Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion
par: Qiao, Yiran, et autres
Publié: (2026)
par: Qiao, Yiran, et autres
Publié: (2026)
Counterfactual Visual Explanation via Causally-Guided Adversarial Steering
par: Qiao, Yiran, et autres
Publié: (2025)
par: Qiao, Yiran, et autres
Publié: (2025)
VIVA: A Benchmark for Vision-Grounded Decision-Making with Human Values
par: Hu, Zhe, et autres
Publié: (2024)
par: Hu, Zhe, et autres
Publié: (2024)
View-consistent Object Removal in Radiance Fields
par: Lu, Yiren, et autres
Publié: (2024)
par: Lu, Yiren, et autres
Publié: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
par: Wang, Xinyu, et autres
Publié: (2025)
par: Wang, Xinyu, et autres
Publié: (2025)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
par: Hu, Zhe, et autres
Publié: (2025)
par: Hu, Zhe, et autres
Publié: (2025)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
par: Xu, Run, et autres
Publié: (2026)
par: Xu, Run, et autres
Publié: (2026)
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
par: Lu, Yiren, et autres
Publié: (2026)
par: Lu, Yiren, et autres
Publié: (2026)
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
par: Shi, Zhengpeng, et autres
Publié: (2025)
par: Shi, Zhengpeng, et autres
Publié: (2025)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
par: He, Jinghan, et autres
Publié: (2024)
par: He, Jinghan, et autres
Publié: (2024)
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
par: Ghosh, Samrajnee, et autres
Publié: (2025)
par: Ghosh, Samrajnee, et autres
Publié: (2025)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
par: Ryan, Yuriel, et autres
Publié: (2025)
par: Ryan, Yuriel, et autres
Publié: (2025)
NeMo: Needle in a Montage for Video-Language Understanding
par: Hu, Zi-Yuan, et autres
Publié: (2025)
par: Hu, Zi-Yuan, et autres
Publié: (2025)
Can Pre-trained Language Models Understand Chinese Humor?
par: Chen, Yuyan, et autres
Publié: (2024)
par: Chen, Yuyan, et autres
Publié: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
par: Zhao, Yi, et autres
Publié: (2026)
par: Zhao, Yi, et autres
Publié: (2026)
Can Language Models Replace Programmers for Coding? REPOCOD Says 'Not Yet'
par: Liang, Shanchao, et autres
Publié: (2024)
par: Liang, Shanchao, et autres
Publié: (2024)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
par: Li, Kaixin, et autres
Publié: (2024)
par: Li, Kaixin, et autres
Publié: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
par: Feng, Xiang, et autres
Publié: (2026)
par: Feng, Xiang, et autres
Publié: (2026)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
par: Chen, Zeren, et autres
Publié: (2023)
par: Chen, Zeren, et autres
Publié: (2023)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
par: Jing, Liqiang, et autres
Publié: (2024)
par: Jing, Liqiang, et autres
Publié: (2024)
Hummus: A Dataset of Humorous Multimodal Metaphor Use
par: Tong, Xiaoyu, et autres
Publié: (2025)
par: Tong, Xiaoyu, et autres
Publié: (2025)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
par: Zhao, Yuwei, et autres
Publié: (2024)
par: Zhao, Yuwei, et autres
Publié: (2024)
Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
par: Zhong, Shanshan, et autres
Publié: (2023)
par: Zhong, Shanshan, et autres
Publié: (2023)
Docopilot: Improving Multimodal Models for Document-Level Understanding
par: Duan, Yuchen, et autres
Publié: (2025)
par: Duan, Yuchen, et autres
Publié: (2025)
DuanzAI: Slang-Enhanced LLM with Prompt for Humor Understanding
par: Rohn, Yesian
Publié: (2024)
par: Rohn, Yesian
Publié: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
par: Wu, Yin, et autres
Publié: (2025)
par: Wu, Yin, et autres
Publié: (2025)
On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation
par: Shang, Wenbo, et autres
Publié: (2026)
par: Shang, Wenbo, et autres
Publié: (2026)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
par: Yin, Yufei, et autres
Publié: (2025)
par: Yin, Yufei, et autres
Publié: (2025)
Investigating the Impact of Rationales for LLMs on Natural Language Understanding
par: Shi, Wenhang, et autres
Publié: (2025)
par: Shi, Wenhang, et autres
Publié: (2025)
Redefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generation
par: Feng, Fu, et autres
Publié: (2024)
par: Feng, Fu, et autres
Publié: (2024)
You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
par: Qin, Zhen, et autres
Publié: (2024)
par: Qin, Zhen, et autres
Publié: (2024)
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
par: Xu, Hongshen, et autres
Publié: (2024)
par: Xu, Hongshen, et autres
Publié: (2024)
Commonality and Individuality! Integrating Humor Commonality with Speaker Individuality for Humor Recognition
par: Zhu, Haohao, et autres
Publié: (2025)
par: Zhu, Haohao, et autres
Publié: (2025)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
par: Wang, Zhaokai, et autres
Publié: (2025)
par: Wang, Zhaokai, et autres
Publié: (2025)
Documents similaires
-
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?
par: Liang, Tuo, et autres
Publié: (2025) -
Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting
par: Lu, Yiren, et autres
Publié: (2025) -
CAUSAL3D: A Comprehensive Benchmark for Causal Learning from Visual Data
par: Liu, Disheng, et autres
Publié: (2025) -
AdvSplat: Adversarial Attacks on Feed-Forward Gaussian Splatting Models
par: Qiao, Yiran, et autres
Publié: (2026) -
DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering
par: Qiao, Yiran, et autres
Publié: (2026)