Toward Cognitive Supersensing in Multimodal Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Boyi, Shen, Yifan, Liu, Yuanzhe, Xu, Yifan, Liu, Jiateng, Li, Xinzhuo, Li, Zhengyuan, Zhu, Jingyuan, Zhong, Yunhan, Lan, Fangzhou, Cao, Jianguo, Rehg, James M., Ji, Heng, Lourentzou, Ismini, Cao, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
von: Yu, Tianjiao, et al.
Veröffentlicht: (2025)
von: Yu, Tianjiao, et al.
Veröffentlicht: (2025)
EgoForge: Goal-Directed Egocentric World Simulator
von: Shen, Yifan, et al.
Veröffentlicht: (2026)
von: Shen, Yifan, et al.
Veröffentlicht: (2026)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
von: Liu, Yuanzhe, et al.
Veröffentlicht: (2026)
von: Liu, Yuanzhe, et al.
Veröffentlicht: (2026)
Evaluating Cognitive Age Alignment in Interactive AI Agents
von: Shen, Yifan, et al.
Veröffentlicht: (2026)
von: Shen, Yifan, et al.
Veröffentlicht: (2026)
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
von: Yu, Tianjiao, et al.
Veröffentlicht: (2026)
von: Yu, Tianjiao, et al.
Veröffentlicht: (2026)
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
von: Li, Xinzhuo, et al.
Veröffentlicht: (2025)
von: Li, Xinzhuo, et al.
Veröffentlicht: (2025)
Cambrian-S: Towards Spatial Supersensing in Video
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
von: Yang, Shusheng, et al.
Veröffentlicht: (2025)
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
von: Cao, Xu, et al.
Veröffentlicht: (2024)
von: Cao, Xu, et al.
Veröffentlicht: (2024)
Commonsense for Zero-Shot Natural Language Video Localization
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
von: Holla, Meghana, et al.
Veröffentlicht: (2023)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
von: Ogunleye, Makanjuola, et al.
Veröffentlicht: (2026)
von: Ogunleye, Makanjuola, et al.
Veröffentlicht: (2026)
mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale
von: Zhou, Xiaona, et al.
Veröffentlicht: (2025)
von: Zhou, Xiaona, et al.
Veröffentlicht: (2025)
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
von: Hao, Haochang, et al.
Veröffentlicht: (2026)
von: Hao, Haochang, et al.
Veröffentlicht: (2026)
Solving Spatial Supersensing Without Spatial Supersensing
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2025)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2025)
MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
von: Tabassum, Afrina, et al.
Veröffentlicht: (2025)
von: Tabassum, Afrina, et al.
Veröffentlicht: (2025)
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
von: Wahed, Muntasir, et al.
Veröffentlicht: (2024)
von: Wahed, Muntasir, et al.
Veröffentlicht: (2024)
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World
von: Wang, Changpeng, et al.
Veröffentlicht: (2026)
von: Wang, Changpeng, et al.
Veröffentlicht: (2026)
Hierarchical Dataset Selection for High-Quality Data Sharing
von: Zhou, Xiaona, et al.
Veröffentlicht: (2025)
von: Zhou, Xiaona, et al.
Veröffentlicht: (2025)
FAIR: Facilitating Artificial Intelligence Resilience in Manufacturing Industrial Internet
von: Zeng, Yingyan, et al.
Veröffentlicht: (2025)
von: Zeng, Yingyan, et al.
Veröffentlicht: (2025)
RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance
von: Venkatesh, Kavana, et al.
Veröffentlicht: (2024)
von: Venkatesh, Kavana, et al.
Veröffentlicht: (2024)
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
von: Shen, Ying, et al.
Veröffentlicht: (2023)
von: Shen, Ying, et al.
Veröffentlicht: (2023)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
von: Shen, Ying, et al.
Veröffentlicht: (2026)
von: Shen, Ying, et al.
Veröffentlicht: (2026)
A Privacy-Preserving Framework for Advertising Personalization Incorporating Federated Learning and Differential Privacy
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
EMCompress: Video-LLMs with Endomorphic Multimodal Compression
von: Fan, Zheyu, et al.
Veröffentlicht: (2025)
von: Fan, Zheyu, et al.
Veröffentlicht: (2025)
Practical Region-level Attack against Segment Anything Models
von: Shen, Yifan, et al.
Veröffentlicht: (2024)
von: Shen, Yifan, et al.
Veröffentlicht: (2024)
Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
von: Wang, Rushi, et al.
Veröffentlicht: (2025)
von: Wang, Rushi, et al.
Veröffentlicht: (2025)
CauSight: Learning to Supersense for Visual Causal Discovery
von: Zhang, Yize, et al.
Veröffentlicht: (2025)
von: Zhang, Yize, et al.
Veröffentlicht: (2025)
Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents
von: Liu, Jiateng, et al.
Veröffentlicht: (2026)
von: Liu, Jiateng, et al.
Veröffentlicht: (2026)
Multimodal Graphene Lubricant Monitoring for IoT ‐Driven Predictive Maintenance
von: Yifan Wang, et al.
Veröffentlicht: (2026)
von: Yifan Wang, et al.
Veröffentlicht: (2026)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
von: Zhou, Xiaona, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaona, et al.
Veröffentlicht: (2026)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
Multimodal Policy Internalization for Conversational Agents
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
Enhancing Speech Large Language Models through Reinforced Behavior Alignment
von: Liu, Yansong, et al.
Veröffentlicht: (2025)
von: Liu, Yansong, et al.
Veröffentlicht: (2025)
Towards Long-horizon Agentic Multimodal Search
von: Du, Yifan, et al.
Veröffentlicht: (2026)
von: Du, Yifan, et al.
Veröffentlicht: (2026)
EVEDIT: Event-based Knowledge Editing with Deductive Editing Boundaries
von: Liu, Jiateng, et al.
Veröffentlicht: (2024)
von: Liu, Jiateng, et al.
Veröffentlicht: (2024)
Liquid directional transport surface applied to the spacecraft fluid management system: Fundamentals and prospect analysis
von: Yifan He, et al.
Veröffentlicht: (2025)
von: Yifan He, et al.
Veröffentlicht: (2025)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
von: Xu, Ruiling, et al.
Veröffentlicht: (2025)
von: Xu, Ruiling, et al.
Veröffentlicht: (2025)
Facile Ester‐based Phase Change Materials Synthesis for Enhanced Energy Storage Toward Battery Thermal Management
von: Long Geng, et al.
Veröffentlicht: (2025)
von: Long Geng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
von: Yu, Tianjiao, et al.
Veröffentlicht: (2025) -
EgoForge: Goal-Directed Egocentric World Simulator
von: Shen, Yifan, et al.
Veröffentlicht: (2026) -
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
von: Shen, Yifan, et al.
Veröffentlicht: (2025) -
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
von: Liu, Yuanzhe, et al.
Veröffentlicht: (2026) -
Evaluating Cognitive Age Alignment in Interactive AI Agents
von: Shen, Yifan, et al.
Veröffentlicht: (2026)