Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Chaudhari, Shravan, Akula, Trilokya, Kim, Yoon, Blake, Tom |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GazeLLM: Multimodal LLMs incorporating Human Visual Attention
por: Rekimoto, Jun
Publicado: (2025)
por: Rekimoto, Jun
Publicado: (2025)
On the Interpretability of Part-Prototype Based Classifiers: A Human Centric Analysis
por: Davoodi, Omid, et al.
Publicado: (2023)
por: Davoodi, Omid, et al.
Publicado: (2023)
MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding
por: Wei, Yuxiang, et al.
Publicado: (2025)
por: Wei, Yuxiang, et al.
Publicado: (2025)
Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception
por: Shen, Junxiao, et al.
Publicado: (2023)
por: Shen, Junxiao, et al.
Publicado: (2023)
TraitSpaces: Towards Interpretable Visual Creativity for Human-AI Co-Creation
por: Luthra, Prerna
Publicado: (2025)
por: Luthra, Prerna
Publicado: (2025)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
por: Natalie, Rosiana, et al.
Publicado: (2025)
por: Natalie, Rosiana, et al.
Publicado: (2025)
LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces
por: Mushkani, Rashid, et al.
Publicado: (2025)
por: Mushkani, Rashid, et al.
Publicado: (2025)
Evaluating Visual Prompts with Eye-Tracking Data for MLLM-Based Human Activity Recognition
por: Choi, Jae Young, et al.
Publicado: (2026)
por: Choi, Jae Young, et al.
Publicado: (2026)
ASAP: Interpretable Analysis and Summarization of AI-generated Image Patterns at Scale
por: Huang, Jinbin, et al.
Publicado: (2024)
por: Huang, Jinbin, et al.
Publicado: (2024)
QPM: Discrete Optimization for Globally Interpretable Image Classification
por: Norrenbrock, Thomas, et al.
Publicado: (2025)
por: Norrenbrock, Thomas, et al.
Publicado: (2025)
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
por: Wu, Hang, et al.
Publicado: (2025)
por: Wu, Hang, et al.
Publicado: (2025)
Posture-Informed Muscular Force Learning for Robust Hand Pressure Estimation
por: Seo, Kyungjin, et al.
Publicado: (2024)
por: Seo, Kyungjin, et al.
Publicado: (2024)
Generalization of CNNs on Relational Reasoning with Bar Charts
por: Cui, Zhenxing, et al.
Publicado: (2025)
por: Cui, Zhenxing, et al.
Publicado: (2025)
Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
por: Jha, Saurav, et al.
Publicado: (2025)
por: Jha, Saurav, et al.
Publicado: (2025)
Agile Deliberation: Concept Deliberation for Subjective Visual Classification
por: Wang, Leijie, et al.
Publicado: (2025)
por: Wang, Leijie, et al.
Publicado: (2025)
VILOD: A Visual Interactive Labeling Tool for Object Detection
por: Holm, Isac
Publicado: (2025)
por: Holm, Isac
Publicado: (2025)
Improving Prototypical Visual Explanations with Reward Reweighing, Reselection, and Retraining
por: Li, Aaron J., et al.
Publicado: (2023)
por: Li, Aaron J., et al.
Publicado: (2023)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
por: Park, Se Jin, et al.
Publicado: (2024)
por: Park, Se Jin, et al.
Publicado: (2024)
AKRMap: Adaptive Kernel Regression for Trustworthy Visualization of Cross-Modal Embeddings
por: Ye, Yilin, et al.
Publicado: (2025)
por: Ye, Yilin, et al.
Publicado: (2025)
Modeling Subjective Urban Perception with Human Gaze
por: Che, Lin, et al.
Publicado: (2026)
por: Che, Lin, et al.
Publicado: (2026)
Looking for a better fit? An Incremental Learning Multimodal Object Referencing Framework adapting to Individual Drivers
por: Gomaa, Amr, et al.
Publicado: (2024)
por: Gomaa, Amr, et al.
Publicado: (2024)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
por: Zhou, Honglu, et al.
Publicado: (2025)
por: Zhou, Honglu, et al.
Publicado: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
por: Wang, Ziwei, et al.
Publicado: (2025)
por: Wang, Ziwei, et al.
Publicado: (2025)
Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
por: Zhang, He, et al.
Publicado: (2025)
por: Zhang, He, et al.
Publicado: (2025)
Generative Augmented Reality: Paradigms, Technologies, and Future Applications
por: Liang, Chen, et al.
Publicado: (2025)
por: Liang, Chen, et al.
Publicado: (2025)
Facial Analysis Systems and Down Syndrome
por: Rondina, Marco, et al.
Publicado: (2025)
por: Rondina, Marco, et al.
Publicado: (2025)
Analysis of the 2024 BraTS Meningioma Radiotherapy Planning Automated Segmentation Challenge
por: LaBella, Dominic, et al.
Publicado: (2024)
por: LaBella, Dominic, et al.
Publicado: (2024)
Context-Awareness and Interpretability of Rare Occurrences for Discovery and Formalization of Critical Failure Modes
por: Polavaram, Sridevi, et al.
Publicado: (2025)
por: Polavaram, Sridevi, et al.
Publicado: (2025)
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
por: Moayeri, Mazda, et al.
Publicado: (2024)
por: Moayeri, Mazda, et al.
Publicado: (2024)
Object Recognition in Human Computer Interaction:- A Comparative Analysis
por: Ranade, Kaushik, et al.
Publicado: (2024)
por: Ranade, Kaushik, et al.
Publicado: (2024)
A Foundational Generative Model for Breast Ultrasound Image Analysis
por: Yu, Haojun, et al.
Publicado: (2025)
por: Yu, Haojun, et al.
Publicado: (2025)
EEG-based Multimodal Representation Learning for Emotion Recognition
por: Yin, Kang, et al.
Publicado: (2024)
por: Yin, Kang, et al.
Publicado: (2024)
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
por: Schoop, Eldon, et al.
Publicado: (2022)
por: Schoop, Eldon, et al.
Publicado: (2022)
Enabling Collaborative Clinical Diagnosis of Infectious Keratitis by Integrating Expert Knowledge and Interpretable Data-driven Intelligence
por: Fang, Zhengqing, et al.
Publicado: (2024)
por: Fang, Zhengqing, et al.
Publicado: (2024)
Magma: A Foundation Model for Multimodal AI Agents
por: Yang, Jianwei, et al.
Publicado: (2025)
por: Yang, Jianwei, et al.
Publicado: (2025)
SASG-DA: Sparse-Aware Semantic-Guided Diffusion Augmentation For Myoelectric Gesture Recognition
por: Liu, Chen, et al.
Publicado: (2025)
por: Liu, Chen, et al.
Publicado: (2025)
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
por: Rajabi, Mohammad Sadra, et al.
Publicado: (2026)
por: Rajabi, Mohammad Sadra, et al.
Publicado: (2026)
Towards a Multimodal Document-grounded Conversational AI System for Education
por: Taneja, Karan, et al.
Publicado: (2025)
por: Taneja, Karan, et al.
Publicado: (2025)
OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
por: Luo, Cheng, et al.
Publicado: (2025)
por: Luo, Cheng, et al.
Publicado: (2025)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
por: Ouyang, Mingyu, et al.
Publicado: (2026)
por: Ouyang, Mingyu, et al.
Publicado: (2026)
Ejemplares similares
-
GazeLLM: Multimodal LLMs incorporating Human Visual Attention
por: Rekimoto, Jun
Publicado: (2025) -
On the Interpretability of Part-Prototype Based Classifiers: A Human Centric Analysis
por: Davoodi, Omid, et al.
Publicado: (2023) -
MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding
por: Wei, Yuxiang, et al.
Publicado: (2025) -
Encode-Store-Retrieve: Augmenting Human Memory through Language-Encoded Egocentric Perception
por: Shen, Junxiao, et al.
Publicado: (2023) -
TraitSpaces: Towards Interpretable Visual Creativity for Human-AI Co-Creation
por: Luthra, Prerna
Publicado: (2025)