MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Saha, Anisha, Suresh, Varsha, Hospedales, Timothy, Demberg, Vera |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
by: Saha, Anisha, et al.
Published: (2026)
by: Saha, Anisha, et al.
Published: (2026)
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
by: Chan, Tsan Tsai, et al.
Published: (2026)
by: Chan, Tsan Tsai, et al.
Published: (2026)
Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
by: Nakai, Toshiki, et al.
Published: (2026)
by: Nakai, Toshiki, et al.
Published: (2026)
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
by: Ahmed, Kareem, et al.
Published: (2026)
by: Ahmed, Kareem, et al.
Published: (2026)
Self-Supervised Multimodal Learning: A Survey
by: Zong, Yongshuo, et al.
Published: (2023)
by: Zong, Yongshuo, et al.
Published: (2023)
Pragmatic Reasoning improves LLM Code Generation
by: Cao, Zhuchen, et al.
Published: (2025)
by: Cao, Zhuchen, et al.
Published: (2025)
RST-LoRA: A Discourse-Aware Low-Rank Adaptation for Long Document Abstractive Summarization
by: Liu, Dongqi, et al.
Published: (2024)
by: Liu, Dongqi, et al.
Published: (2024)
ChatGPT vs Human-authored Text: Insights into Controllable Text Summarization and Sentence Style Transfer
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
by: Damani, Mehul, et al.
Published: (2025)
by: Damani, Mehul, et al.
Published: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
by: Dang, Vy Tuong, et al.
Published: (2025)
by: Dang, Vy Tuong, et al.
Published: (2025)
SemIRNet: A Semantic Irony Recognition Network for Multimodal Sarcasm Detection
by: Zhou, Jingxuan, et al.
Published: (2025)
by: Zhou, Jingxuan, et al.
Published: (2025)
Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction
by: Saha, Anisha, et al.
Published: (2025)
by: Saha, Anisha, et al.
Published: (2025)
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
by: Zhao, Zehua, et al.
Published: (2025)
by: Zhao, Zehua, et al.
Published: (2025)
On Sarcasm Detection with OpenAI GPT-based Models
by: Gole, Montgomery, et al.
Published: (2023)
by: Gole, Montgomery, et al.
Published: (2023)
Sarcasm Detection in Tweets with BERT and GloVe Embeddings
by: Khatri, Akshay, et al.
Published: (2020)
by: Khatri, Akshay, et al.
Published: (2020)
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization
by: Liu, Dongqi, et al.
Published: (2023)
by: Liu, Dongqi, et al.
Published: (2023)
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues
by: Suresh, Varsha, et al.
Published: (2025)
by: Suresh, Varsha, et al.
Published: (2025)
Modeling Turn-Taking with Semantically Informed Gestures
by: Suresh, Varsha, et al.
Published: (2025)
by: Suresh, Varsha, et al.
Published: (2025)
RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Sarcasm Detection on Reddit Using Classical Machine Learning and Feature Engineering
by: Karmaker, Subrata
Published: (2025)
by: Karmaker, Subrata
Published: (2025)
SciNews: From Scholarly Complexities to Public Narratives -- A Dataset for Scientific News Report Generation
by: Liu, Dongqi, et al.
Published: (2024)
by: Liu, Dongqi, et al.
Published: (2024)
MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs
by: Roccabruna, Gabriel, et al.
Published: (2026)
by: Roccabruna, Gabriel, et al.
Published: (2026)
Synthetic Data Augmentation for Cross-domain Implicit Discourse Relation Recognition
by: Yung, Frances, et al.
Published: (2025)
by: Yung, Frances, et al.
Published: (2025)
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning
by: Cai, Zikui, et al.
Published: (2025)
by: Cai, Zikui, et al.
Published: (2025)
Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage
by: He, Ziyi, et al.
Published: (2026)
by: He, Ziyi, et al.
Published: (2026)
Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
Explanatory Summarization with Discourse-Driven Planning
by: Liu, Dongqi, et al.
Published: (2025)
by: Liu, Dongqi, et al.
Published: (2025)
Compositional preference models for aligning LMs
by: Go, Dongyoung, et al.
Published: (2023)
by: Go, Dongyoung, et al.
Published: (2023)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
by: Nath, Vaskar, et al.
Published: (2025)
by: Nath, Vaskar, et al.
Published: (2025)
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding
by: Cheshmi, Seyyed Saeid, et al.
Published: (2026)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2026)
Enabling Approximate Joint Sampling in Diffusion LMs
by: Bansal, Parikshit, et al.
Published: (2025)
by: Bansal, Parikshit, et al.
Published: (2025)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
by: Potamitis, Nearchos, et al.
Published: (2025)
by: Potamitis, Nearchos, et al.
Published: (2025)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
When Do "More Contexts" Help with Sarcasm Recognition?
by: Nimase, Ojas, et al.
Published: (2024)
by: Nimase, Ojas, et al.
Published: (2024)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
by: Oh, Jungwoo, et al.
Published: (2026)
by: Oh, Jungwoo, et al.
Published: (2026)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
by: Ni, Ruikang, et al.
Published: (2024)
by: Ni, Ruikang, et al.
Published: (2024)
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024)
by: Seto, Skyler, et al.
Published: (2024)
LLMs as Assessors: Right for the Right Reason?
by: Saha, Sourav, et al.
Published: (2026)
by: Saha, Sourav, et al.
Published: (2026)
Similar Items
-
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
by: Saha, Anisha, et al.
Published: (2026) -
System-Mediated Attention Imbalances Make Vision-Language Models Say Yes
by: Chan, Tsan Tsai, et al.
Published: (2026) -
Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models
by: Nakai, Toshiki, et al.
Published: (2026) -
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
by: Ahmed, Kareem, et al.
Published: (2026) -
Self-Supervised Multimodal Learning: A Survey
by: Zong, Yongshuo, et al.
Published: (2023)