Emotional Theory of Mind: Bridging Fast Visual Processing with Slow Linguistic Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Etesam, Yasaman, Yalçın, Özge Nilay, Zhang, Chuxuan, Lim, Angelica |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contextual Emotion Recognition using Large Vision Language Models
by: Etesam, Yasaman, et al.
Published: (2024)
by: Etesam, Yasaman, et al.
Published: (2024)
Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents
by: Lim, Angelica, et al.
Published: (2026)
by: Lim, Angelica, et al.
Published: (2026)
React to This (RTT): A Nonverbal Turing Test for Embodied AI
by: Zhang, Chuxuan, et al.
Published: (2025)
by: Zhang, Chuxuan, et al.
Published: (2025)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
by: Luo, Meng, et al.
Published: (2026)
by: Luo, Meng, et al.
Published: (2026)
Learning to Think Fast and Slow for Visual Language Models
by: Lin, Chenyu, et al.
Published: (2025)
by: Lin, Chenyu, et al.
Published: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
by: Chen, Zhenghao, et al.
Published: (2026)
by: Chen, Zhenghao, et al.
Published: (2026)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
by: Zeng, Haijin, et al.
Published: (2025)
by: Zeng, Haijin, et al.
Published: (2025)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024)
by: Takmaz, Ece, et al.
Published: (2024)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
by: Wei, Yana, et al.
Published: (2025)
by: Wei, Yana, et al.
Published: (2025)
SignAgent: Agentic LLMs for Linguistically-Grounded Sign Language Annotation and Dataset Curation
by: Cory, Oliver, et al.
Published: (2026)
by: Cory, Oliver, et al.
Published: (2026)
Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning
by: Yang, Jingru, et al.
Published: (2024)
by: Yang, Jingru, et al.
Published: (2024)
FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models
by: Yang, Yiming, et al.
Published: (2025)
by: Yang, Yiming, et al.
Published: (2025)
Gender Bias in Emotion Recognition by Large Language Models
by: Herbert, Maureen, et al.
Published: (2025)
by: Herbert, Maureen, et al.
Published: (2025)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
by: Chen, Weixing, et al.
Published: (2025)
by: Chen, Weixing, et al.
Published: (2025)
Bridge the Gap Between Visual and Linguistic Comprehension for Generalized Zero-shot Semantic Segmentation
by: Guo, Xiaoqing, et al.
Published: (2025)
by: Guo, Xiaoqing, et al.
Published: (2025)
CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process
by: Akbari, Arman, et al.
Published: (2025)
by: Akbari, Arman, et al.
Published: (2025)
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
by: Yin, Tianwei, et al.
Published: (2024)
by: Yin, Tianwei, et al.
Published: (2024)
Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
by: Burkanova, Bermet, et al.
Published: (2025)
by: Burkanova, Bermet, et al.
Published: (2025)
Enhancing Visual Place Recognition via Fast and Slow Adaptive Biasing in Event Cameras
by: Nair, Gokul B., et al.
Published: (2024)
by: Nair, Gokul B., et al.
Published: (2024)
VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory
by: Wang, Shaoan, et al.
Published: (2026)
by: Wang, Shaoan, et al.
Published: (2026)
Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation
by: Sincan, Ozge Mercanoglu, et al.
Published: (2025)
by: Sincan, Ozge Mercanoglu, et al.
Published: (2025)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
by: Xiao, Wenyi, et al.
Published: (2025)
by: Xiao, Wenyi, et al.
Published: (2025)
Bridging Text and Video Generation: A Survey
by: Kumar, Nilay, et al.
Published: (2025)
by: Kumar, Nilay, et al.
Published: (2025)
Fast-Slow Efficient Training for Multimodal Large Language Models via Visual Token Pruning
by: Zhang, Dingkun, et al.
Published: (2026)
by: Zhang, Dingkun, et al.
Published: (2026)
VAEER: Visual Attention-Inspired Emotion Elicitation Reasoning
by: Man, Fanhang, et al.
Published: (2025)
by: Man, Fanhang, et al.
Published: (2025)
Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
by: Jia, Guoli, et al.
Published: (2025)
by: Jia, Guoli, et al.
Published: (2025)
Slot-VLM: SlowFast Slots for Video-Language Modeling
by: Xu, Jiaqi, et al.
Published: (2024)
by: Xu, Jiaqi, et al.
Published: (2024)
AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
by: Zhang, Ruifei, et al.
Published: (2025)
by: Zhang, Ruifei, et al.
Published: (2025)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
by: Nazi, Zabir Al, et al.
Published: (2025)
by: Nazi, Zabir Al, et al.
Published: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
by: Li, Chenxuan, et al.
Published: (2024)
by: Li, Chenxuan, et al.
Published: (2024)
Bridging the Intent Gap: Knowledge-Enhanced Visual Generation
by: Cheng, Yi, et al.
Published: (2024)
by: Cheng, Yi, et al.
Published: (2024)
Fast-Slow Test-Time Adaptation for Online Vision-and-Language Navigation
by: Gao, Junyu, et al.
Published: (2023)
by: Gao, Junyu, et al.
Published: (2023)
Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
by: Schusterbauer, Johannes, et al.
Published: (2026)
by: Schusterbauer, Johannes, et al.
Published: (2026)
iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
by: Mehrotra, Sarthak, et al.
Published: (2025)
by: Mehrotra, Sarthak, et al.
Published: (2025)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
by: Shi, Min, et al.
Published: (2025)
by: Shi, Min, et al.
Published: (2025)
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
by: Pande, Nilay, et al.
Published: (2025)
by: Pande, Nilay, et al.
Published: (2025)
Language-Conditioned Robotic Manipulation with Fast and Slow Thinking
by: Zhu, Minjie, et al.
Published: (2024)
by: Zhu, Minjie, et al.
Published: (2024)
Flow-CDNet: A Novel Network for Detecting Both Slow and Fast Changes in Bitemporal Images
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
HyperZ$\cdot$Z$\cdot$W Operator Connects Slow-Fast Networks for Full Context Interaction
by: Zhang, Harvie
Published: (2024)
by: Zhang, Harvie
Published: (2024)
Similar Items
-
Contextual Emotion Recognition using Large Vision Language Models
by: Etesam, Yasaman, et al.
Published: (2024) -
Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents
by: Lim, Angelica, et al.
Published: (2026) -
React to This (RTT): A Nonverbal Turing Test for Embodied AI
by: Zhang, Chuxuan, et al.
Published: (2025) -
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
by: Luo, Meng, et al.
Published: (2026) -
Learning to Think Fast and Slow for Visual Language Models
by: Lin, Chenyu, et al.
Published: (2025)