Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chaturvedi, Isha, Nair, Anjana, Li, Yushen, Kumar, Adhitya Rajendra, Zhu, Kevin, Dev, Sunishchal, Panda, Ashwinee, Sharma, Vasu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
von: Egbuna, Nathan, et al.
Veröffentlicht: (2025)
von: Egbuna, Nathan, et al.
Veröffentlicht: (2025)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
von: Swaroop, Anand, et al.
Veröffentlicht: (2025)
von: Swaroop, Anand, et al.
Veröffentlicht: (2025)
The Geometry of Harmfulness in LLMs through Subconcept Probing
von: Shah, McNair, et al.
Veröffentlicht: (2025)
von: Shah, McNair, et al.
Veröffentlicht: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
von: Agrawal, Shriyansh, et al.
Veröffentlicht: (2025)
von: Agrawal, Shriyansh, et al.
Veröffentlicht: (2025)
Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
von: Su, Chris, et al.
Veröffentlicht: (2025)
von: Su, Chris, et al.
Veröffentlicht: (2025)
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
von: Rana, Manik, et al.
Veröffentlicht: (2025)
von: Rana, Manik, et al.
Veröffentlicht: (2025)
Broken Chains: The Cost of Incomplete Reasoning in LLMs
von: Su, Ian, et al.
Veröffentlicht: (2026)
von: Su, Ian, et al.
Veröffentlicht: (2026)
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
von: Mao, Nathan, et al.
Veröffentlicht: (2026)
von: Mao, Nathan, et al.
Veröffentlicht: (2026)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
von: Patel, Dev, et al.
Veröffentlicht: (2025)
von: Patel, Dev, et al.
Veröffentlicht: (2025)
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
von: Chang, Vincent, et al.
Veröffentlicht: (2025)
von: Chang, Vincent, et al.
Veröffentlicht: (2025)
PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
von: Vuddanti, Sri Vatsa, et al.
Veröffentlicht: (2025)
von: Vuddanti, Sri Vatsa, et al.
Veröffentlicht: (2025)
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
von: Lu, Leo, et al.
Veröffentlicht: (2025)
von: Lu, Leo, et al.
Veröffentlicht: (2025)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
von: Arturi, Daniel Aarao Reis, et al.
Veröffentlicht: (2025)
von: Arturi, Daniel Aarao Reis, et al.
Veröffentlicht: (2025)
SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning
von: Aluru, Aayush, et al.
Veröffentlicht: (2025)
von: Aluru, Aayush, et al.
Veröffentlicht: (2025)
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
von: Zhang, Juzheng, et al.
Veröffentlicht: (2025)
von: Zhang, Juzheng, et al.
Veröffentlicht: (2025)
CA-BED: Conversation-Aware Bayesian Experimental Design
von: Arnould, Daniel, et al.
Veröffentlicht: (2026)
von: Arnould, Daniel, et al.
Veröffentlicht: (2026)
From Refusal Tokens to Refusal Control: Discovering and Steering Category-Specific Refusal Directions
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
von: Alagharu, Rishab, et al.
Veröffentlicht: (2026)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners
von: Jiang, Bowen, et al.
Veröffentlicht: (2024)
von: Jiang, Bowen, et al.
Veröffentlicht: (2024)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
von: Panda, Ashwinee, et al.
Veröffentlicht: (2022)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2022)
FairReason: Balancing Reasoning and Social Bias in MLLMs
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
von: Pan, Zhenyu, et al.
Veröffentlicht: (2025)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
von: He, Jiahang, et al.
Veröffentlicht: (2025)
von: He, Jiahang, et al.
Veröffentlicht: (2025)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
von: Helu, Zhi, et al.
Veröffentlicht: (2025)
von: Helu, Zhi, et al.
Veröffentlicht: (2025)
AuditCopilot: Leveraging LLMs for Fraud Detection in Double-Entry Bookkeeping
von: Kadir, Md Abdul, et al.
Veröffentlicht: (2025)
von: Kadir, Md Abdul, et al.
Veröffentlicht: (2025)
Privacy Auditing of Large Language Models
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2025)
GamiBench: Evaluating Spatial Reasoning and 2D-to-3D Planning Capabilities of MLLMs with Origami Folding Tasks
von: Spencer, Ryan, et al.
Veröffentlicht: (2025)
von: Spencer, Ryan, et al.
Veröffentlicht: (2025)
InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
Up to 36x Speedup: Mask-based Parallel Inference Paradigm for Key Information Extraction in MLLMs
von: Wang, Xinzhong, et al.
Veröffentlicht: (2026)
von: Wang, Xinzhong, et al.
Veröffentlicht: (2026)
ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2025)
von: Xu, Zhengzhuo, et al.
Veröffentlicht: (2025)
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
von: Ye, Hengwei, et al.
Veröffentlicht: (2026)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
von: Xiong, Lang, et al.
Veröffentlicht: (2025)
von: Xiong, Lang, et al.
Veröffentlicht: (2025)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
von: Begin, James, et al.
Veröffentlicht: (2025)
von: Begin, James, et al.
Veröffentlicht: (2025)
The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus
von: Jindal, Ishan, et al.
Veröffentlicht: (2026)
von: Jindal, Ishan, et al.
Veröffentlicht: (2026)
Teach LLMs to Phish: Stealing Private Information from Language Models
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
von: Batra, Shourya, et al.
Veröffentlicht: (2025) -
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
von: Egbuna, Nathan, et al.
Veröffentlicht: (2025) -
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
von: Swaroop, Anand, et al.
Veröffentlicht: (2025) -
The Geometry of Harmfulness in LLMs through Subconcept Probing
von: Shah, McNair, et al.
Veröffentlicht: (2025) -
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
von: Thomas, Rohan Subramanian, et al.
Veröffentlicht: (2026)