MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Haochen, Kong, Yuyao, Xu, Yongxiu, Gou, Gaopeng, Xu, Hongbo, Wang, Yubin, Zhang, Haoliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts
von: Lin, Xinkui, et al.
Veröffentlicht: (2025)
von: Lin, Xinkui, et al.
Veröffentlicht: (2025)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026)
MMSD-Net: Towards Multi-modal Stuttering Detection
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
von: Xu, Jingning, et al.
Veröffentlicht: (2026)
von: Xu, Jingning, et al.
Veröffentlicht: (2026)
T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
von: Li, Yili, et al.
Veröffentlicht: (2025)
von: Li, Yili, et al.
Veröffentlicht: (2025)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
von: Chen, Shuang, et al.
Veröffentlicht: (2026)
LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment
von: Hao, Fangyu, et al.
Veröffentlicht: (2026)
von: Hao, Fangyu, et al.
Veröffentlicht: (2026)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
von: Peng, Zixiang, et al.
Veröffentlicht: (2026)
von: Peng, Zixiang, et al.
Veröffentlicht: (2026)
DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
von: Zhao, Kangran, et al.
Veröffentlicht: (2025)
von: Zhao, Kangran, et al.
Veröffentlicht: (2025)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
von: Xu, Shuolin, et al.
Veröffentlicht: (2025)
Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
von: Sun, Shengyang, et al.
Veröffentlicht: (2024)
von: Sun, Shengyang, et al.
Veröffentlicht: (2024)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
von: Liu, Hui, et al.
Veröffentlicht: (2023)
von: Liu, Hui, et al.
Veröffentlicht: (2023)
Kandinsky 3.0 Technical Report
von: Arkhipkin, Vladimir, et al.
Veröffentlicht: (2023)
von: Arkhipkin, Vladimir, et al.
Veröffentlicht: (2023)
HarmonyIQA: Pioneering Benchmark and Model for Image Harmonization Quality Assessment
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
von: Chen, Junjie, et al.
Veröffentlicht: (2025)
"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
von: Skoularikis, Anastasios, et al.
Veröffentlicht: (2025)
von: Skoularikis, Anastasios, et al.
Veröffentlicht: (2025)
FakeBench: Probing Explainable Fake Image Detection via Large Multimodal Models
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
MFQE 2.0: A New Approach for Multi-frame Quality Enhancement on Compressed Video
von: Xing, Qunliang, et al.
Veröffentlicht: (2019)
von: Xing, Qunliang, et al.
Veröffentlicht: (2019)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
von: Wang, Bing, et al.
Veröffentlicht: (2024)
von: Wang, Bing, et al.
Veröffentlicht: (2024)
Detached and Interactive Multimodal Learning
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
von: Ikuta, Hikaru, et al.
Veröffentlicht: (2024)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
von: Guo, Zile, et al.
Veröffentlicht: (2026)
von: Guo, Zile, et al.
Veröffentlicht: (2026)
DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning
von: Lu, Jialang, et al.
Veröffentlicht: (2025)
von: Lu, Jialang, et al.
Veröffentlicht: (2025)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts
von: Lin, Xinkui, et al.
Veröffentlicht: (2025) -
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
von: Wang, Zhenyu, et al.
Veröffentlicht: (2026) -
MMSD-Net: Towards Multi-modal Stuttering Detection
von: Nie, Liangyu, et al.
Veröffentlicht: (2024) -
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
von: Kong, Chenqi, et al.
Veröffentlicht: (2023) -
PDA: Text-Augmented Defense Framework for Robust Vision-Language Models against Adversarial Image Attacks
von: Xu, Jingning, et al.
Veröffentlicht: (2026)