Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Qianqi, Li, Hongquan, Jiang, Shan, Zhao, Yang, Guan, Xinze, Kuo, Ching-Chen, Wang, Xin Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
by: Yan, Qianqi, et al.
Published: (2025)
by: Yan, Qianqi, et al.
Published: (2025)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
Hidden in Plain Sight: Exploring Chat History Tampering in Interactive Language Models
by: Wei, Cheng'an, et al.
Published: (2024)
by: Wei, Cheng'an, et al.
Published: (2024)
Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA
by: Yan, Qianqi, et al.
Published: (2024)
by: Yan, Qianqi, et al.
Published: (2024)
Hidden in Plain Sight: Visual-to-Symbolic Analytical Solution Inference from Field Visualizations
by: Li, Pengze, et al.
Published: (2026)
by: Li, Pengze, et al.
Published: (2026)
The Robustness of Differentiable Causal Discovery in Misspecified Scenarios
by: Yi, Huiyang, et al.
Published: (2025)
by: Yi, Huiyang, et al.
Published: (2025)
The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
by: Harvey, Thomas R.
Published: (2025)
by: Harvey, Thomas R.
Published: (2025)
Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models
by: Hartman, Max, et al.
Published: (2026)
by: Hartman, Max, et al.
Published: (2026)
Underspecified Human Decision Experiments Considered Harmful
by: Hullman, Jessica, et al.
Published: (2024)
by: Hullman, Jessica, et al.
Published: (2024)
Hidden in Plain Sight: Undetectable Adversarial Bias Attacks on Vulnerable Patient Populations
by: Kulkarni, Pranav, et al.
Published: (2024)
by: Kulkarni, Pranav, et al.
Published: (2024)
CausalCompass: Evaluating the Robustness of Time-Series Causal Discovery in Misspecified Scenarios
by: Yi, Huiyang, et al.
Published: (2026)
by: Yi, Huiyang, et al.
Published: (2026)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
by: Huang, Zhongzhen, et al.
Published: (2025)
by: Huang, Zhongzhen, et al.
Published: (2025)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
by: Lu, Yu-An, et al.
Published: (2026)
by: Lu, Yu-An, et al.
Published: (2026)
Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference
by: Chen, Depeng, et al.
Published: (2024)
by: Chen, Depeng, et al.
Published: (2024)
MemoSight: Unifying Context Compression and Multi Token Prediction for Reasoning Acceleration
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
Honeyfile Camouflage: Hiding Fake Files in Plain Sight
by: Timmer, Roelien C., et al.
Published: (2024)
by: Timmer, Roelien C., et al.
Published: (2024)
From Attribution to Abstention: Training-Free Attention-Based Auditing for Clinical Summarization
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
by: Geng, Jianing, et al.
Published: (2025)
by: Geng, Jianing, et al.
Published: (2025)
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs
by: Shu, Yubo, et al.
Published: (2025)
by: Shu, Yubo, et al.
Published: (2025)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
Statistical Inference for Misspecified Contextual Bandits
by: Guo, Yongyi, et al.
Published: (2025)
by: Guo, Yongyi, et al.
Published: (2025)
ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
by: Zhang, Zhen, et al.
Published: (2025)
by: Zhang, Zhen, et al.
Published: (2025)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
by: Jiang, Ruixiang, et al.
Published: (2025)
by: Jiang, Ruixiang, et al.
Published: (2025)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
by: Yang, Tianyu, et al.
Published: (2026)
by: Yang, Tianyu, et al.
Published: (2026)
A Discordance-Aware Multimodal Framework with Multi-Agent Clinical Reasoning
by: Ahadian, Pegah, et al.
Published: (2026)
by: Ahadian, Pegah, et al.
Published: (2026)
Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
by: Hoover, Benjamin, et al.
Published: (2023)
by: Hoover, Benjamin, et al.
Published: (2023)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
by: Zhou, Ruiwen, et al.
Published: (2024)
by: Zhou, Ruiwen, et al.
Published: (2024)
Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
by: Li, Junxian, et al.
Published: (2025)
by: Li, Junxian, et al.
Published: (2025)
Multi-Scenario Reasoning: Unlocking Cognitive Autonomy in Humanoid Robots for Multimodal Understanding
by: Wang, Libo
Published: (2024)
by: Wang, Libo
Published: (2024)
Co-Sight: Enhancing LLM-Based Agents via Conflict-Aware Meta-Verification and Trustworthy Reasoning with Structured Facts
by: Zhang, Hongwei, et al.
Published: (2025)
by: Zhang, Hongwei, et al.
Published: (2025)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
by: Xu, Zelin, et al.
Published: (2026)
by: Xu, Zelin, et al.
Published: (2026)
Seeing with You: Perception-Reasoning Coevolution for Multimodal Reasoning
by: Miao, Ziqi, et al.
Published: (2026)
by: Miao, Ziqi, et al.
Published: (2026)
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
by: Luo, Shixian, et al.
Published: (2025)
by: Luo, Shixian, et al.
Published: (2025)
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
by: Wang, Ben, et al.
Published: (2026)
by: Wang, Ben, et al.
Published: (2026)
The Hidden Threat in Plain Text: Attacking RAG Data Loaders
by: Castagnaro, Alberto, et al.
Published: (2025)
by: Castagnaro, Alberto, et al.
Published: (2025)
Similar Items
-
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
by: Yan, Qianqi, et al.
Published: (2025) -
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024) -
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
by: Yan, Qianqi, et al.
Published: (2026) -
Hidden in Plain Sight: Exploring Chat History Tampering in Interactive Language Models
by: Wei, Cheng'an, et al.
Published: (2024) -
Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA
by: Yan, Qianqi, et al.
Published: (2024)