HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Zheng, Zheng, Ruobing, Wang, Yabing, Li, Tianqi, Yuan, Yi, Chen, Jingdong, Wang, Le |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Versatile Multimodal Controls for Expressive Talking Human Animation
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
by: Zheng, Ruobing, et al.
Published: (2026)
by: Zheng, Ruobing, et al.
Published: (2026)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation
by: Wang, Liping, et al.
Published: (2026)
by: Wang, Liping, et al.
Published: (2026)
LokiTalk: Learning Fine-Grained and Generalizable Correspondences to Enhance NeRF-based Talking Head Synthesis
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
by: Wang, Yabing, et al.
Published: (2025)
by: Wang, Yabing, et al.
Published: (2025)
GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing
by: Sun, Qigan, et al.
Published: (2026)
by: Sun, Qigan, et al.
Published: (2026)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
by: Wang, Kuo, et al.
Published: (2025)
by: Wang, Kuo, et al.
Published: (2025)
Incentivizing Cardiologist-Like Reasoning in MLLMs for Interpretable Echocardiographic Diagnosis
by: Qin, Yi, et al.
Published: (2026)
by: Qin, Yi, et al.
Published: (2026)
Interpreting and Mitigating Hallucination in MLLMs through Multi-agent Debate
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
Learning Dynamic Tetrahedra for High-Quality Talking Head Synthesis
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
by: Wang, Jianghui, et al.
Published: (2023)
by: Wang, Jianghui, et al.
Published: (2023)
Animate-X: Universal Character Image Animation with Enhanced Motion Representation
by: Tan, Shuai, et al.
Published: (2024)
by: Tan, Shuai, et al.
Published: (2024)
GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Perception-Aware Multimodal Spatial Reasoning from Monocular Images
by: Cheng, Yanchun, et al.
Published: (2026)
by: Cheng, Yanchun, et al.
Published: (2026)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
by: Zhu, Fangrui, et al.
Published: (2025)
by: Zhu, Fangrui, et al.
Published: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
by: Zhang, Shan, et al.
Published: (2025)
by: Zhang, Shan, et al.
Published: (2025)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
by: Chen, Zhenghao, et al.
Published: (2026)
by: Chen, Zhenghao, et al.
Published: (2026)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
by: Wang, Haicheng, et al.
Published: (2026)
by: Wang, Haicheng, et al.
Published: (2026)
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
by: Liu, Ruiheng, et al.
Published: (2026)
by: Liu, Ruiheng, et al.
Published: (2026)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
by: Hao, Yunzhuo, et al.
Published: (2025)
by: Hao, Yunzhuo, et al.
Published: (2025)
SASP: Strip-Aware Spatial Perception for Fine-Grained Bird Image Classification
by: Wang, Zheng
Published: (2025)
by: Wang, Zheng
Published: (2025)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
by: Liu, Huan, et al.
Published: (2024)
by: Liu, Huan, et al.
Published: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
by: Li, Wen, et al.
Published: (2024)
by: Li, Wen, et al.
Published: (2024)
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025)
by: Zou, Xin, et al.
Published: (2025)
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
by: Li, Xudong, et al.
Published: (2025)
by: Li, Xudong, et al.
Published: (2025)
Collaborative Position Reasoning Network for Referring Image Segmentation
by: Cao, Jianjian, et al.
Published: (2024)
by: Cao, Jianjian, et al.
Published: (2024)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
Similar Items
-
Versatile Multimodal Controls for Expressive Talking Human Animation
by: Qin, Zheng, et al.
Published: (2025) -
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
by: Zheng, Ruobing, et al.
Published: (2026) -
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025) -
RSRNav: Reasoning Spatial Relationship for Image-Goal Navigation
by: Qin, Zheng, et al.
Published: (2025) -
Embracing Aleatoric Uncertainty: Generating Diverse 3D Human Motion
by: Qin, Zheng, et al.
Published: (2025)