Gespeichert in:
| Hauptverfasser: | Luo, Sha, Prabhu, Yogesh, Ossowski, Timothy, Chen, Kaiping, Hu, Junjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.03369 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prompting Large Vision-Language Models for Compositional Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
von: Chen, Wei, et al.
Veröffentlicht: (2025)
von: Chen, Wei, et al.
Veröffentlicht: (2025)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
von: Zhang, Zory, et al.
Veröffentlicht: (2025)
von: Zhang, Zory, et al.
Veröffentlicht: (2025)
OLIVE: Object Level In-Context Visual Embeddings
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)
Video-ToC: Video Tree-of-Cue Reasoning
von: Tan, Qizhong, et al.
Veröffentlicht: (2026)
von: Tan, Qizhong, et al.
Veröffentlicht: (2026)
Learning Multimodal Cues of Children's Uncertainty
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
von: Kim, Youngmin, et al.
Veröffentlicht: (2025)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
von: Hinojosa, Carlos, et al.
Veröffentlicht: (2026)
von: Hinojosa, Carlos, et al.
Veröffentlicht: (2026)
Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization
von: Islam, Md Moinul, et al.
Veröffentlicht: (2025)
von: Islam, Md Moinul, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
von: Feng, X., et al.
Veröffentlicht: (2024)
von: Feng, X., et al.
Veröffentlicht: (2024)
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
von: Hosseini, Parsa, et al.
Veröffentlicht: (2025)
von: Hosseini, Parsa, et al.
Veröffentlicht: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
von: Liang, Xin, et al.
Veröffentlicht: (2025)
von: Liang, Xin, et al.
Veröffentlicht: (2025)
CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
von: Yu, Yating, et al.
Veröffentlicht: (2025)
von: Yu, Yating, et al.
Veröffentlicht: (2025)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
von: Yang, Luyu, et al.
Veröffentlicht: (2026)
von: Yang, Luyu, et al.
Veröffentlicht: (2026)
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
von: Wang, Andong, et al.
Veröffentlicht: (2024)
von: Wang, Andong, et al.
Veröffentlicht: (2024)
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
von: Sarto, Sara, et al.
Veröffentlicht: (2024)
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
von: Li, Siqi, et al.
Veröffentlicht: (2025)
von: Li, Siqi, et al.
Veröffentlicht: (2025)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
von: Chen, Kaijie, et al.
Veröffentlicht: (2025)
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2025)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Prompting Large Vision-Language Models for Compositional Reasoning
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024) -
Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
von: Chen, Wei, et al.
Veröffentlicht: (2025) -
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
von: Qi, Yukun, et al.
Veröffentlicht: (2026) -
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
von: Zhang, Zory, et al.
Veröffentlicht: (2025) -
OLIVE: Object Level In-Context Visual Embeddings
von: Ossowski, Timothy, et al.
Veröffentlicht: (2024)