Can Large Vision-Language Models Understand Multimodal Sarcasm?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinyu, Zhang, Yue, Jing, Liqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
von: Chang, Yue, et al.
Veröffentlicht: (2024)
von: Chang, Yue, et al.
Veröffentlicht: (2024)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
Do Multimodal Large Language Models Understand Welding?
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
PUMGPT: A Large Vision-Language Model for Product Understanding
von: Xue, Wei, et al.
Veröffentlicht: (2023)
von: Xue, Wei, et al.
Veröffentlicht: (2023)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
Are Large Vision Language Models Good Game Players?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Toward Interactive Regional Understanding in Vision-Large Language Models
von: Lee, Jungbeom, et al.
Veröffentlicht: (2024)
von: Lee, Jungbeom, et al.
Veröffentlicht: (2024)
Can Vision Language Models Understand Mimed Actions?
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
von: Cho, Hyundong, et al.
Veröffentlicht: (2025)
GDCNet: Generative Discrepancy Comparison Network for Multimodal Sarcasm Detection
von: Zhang, Shuguang, et al.
Veröffentlicht: (2026)
von: Zhang, Shuguang, et al.
Veröffentlicht: (2026)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
von: Li, Yun, et al.
Veröffentlicht: (2025)
von: Li, Yun, et al.
Veröffentlicht: (2025)
Can We Predict Performance of Large Models across Vision-Language Tasks?
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Video Understanding with Large Language Models: A Survey
von: Tang, Yolo Y., et al.
Veröffentlicht: (2023)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2023)
SemIRNet: A Semantic Irony Recognition Network for Multimodal Sarcasm Detection
von: Zhou, Jingxuan, et al.
Veröffentlicht: (2025)
von: Zhou, Jingxuan, et al.
Veröffentlicht: (2025)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
BLINK: Multimodal Large Language Models Can See but Not Perceive
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
von: Xing, Shangyu, et al.
Veröffentlicht: (2024)
von: Xing, Shangyu, et al.
Veröffentlicht: (2024)
M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Can Vision-Language Models Evaluate Handwritten Math?
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
Can Vision-Language Models Solve the Shell Game?
von: Liu, Tiedong, et al.
Veröffentlicht: (2026)
von: Liu, Tiedong, et al.
Veröffentlicht: (2026)
Multimodal Abstractive Summarization of Instructional Videos with Vision-Language Models
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
von: Nazir, Maham, et al.
Veröffentlicht: (2026)
Vision-Language Models Do Not Understand Negation
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2024)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
von: Lin, Junyan, et al.
Veröffentlicht: (2026)
von: Lin, Junyan, et al.
Veröffentlicht: (2026)
Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions
von: Hu, Zhe, et al.
Veröffentlicht: (2024)
von: Hu, Zhe, et al.
Veröffentlicht: (2024)
Do Vision-Language Models Really Understand Visual Language?
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
von: Hu, Jinyi, et al.
Veröffentlicht: (2023)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
von: Chang, Yue, et al.
Veröffentlicht: (2024) -
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
von: Jing, Liqiang, et al.
Veröffentlicht: (2024) -
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025) -
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024) -
Do Multimodal Large Language Models Understand Welding?
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)