EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Seoh, Ronald, Goldwasser, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
by: Lu, Xingyuan, et al.
Published: (2025)
by: Lu, Xingyuan, et al.
Published: (2025)
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
by: Wu, Chengfei, et al.
Published: (2025)
by: Wu, Chengfei, et al.
Published: (2025)
Internalized Reasoning for Long-Context Visual Document Understanding
by: Veselka, Austin
Published: (2026)
by: Veselka, Austin
Published: (2026)
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
by: Xie, Hongxia, et al.
Published: (2024)
by: Xie, Hongxia, et al.
Published: (2024)
EmoCAM: Toward Understanding What Drives CNN-based Emotion Recognition
by: Doulfoukar, Youssef, et al.
Published: (2024)
by: Doulfoukar, Youssef, et al.
Published: (2024)
Retrieving Counterfactuals Improves Visual In-Context Learning
by: Xiong, Guangzhi, et al.
Published: (2026)
by: Xiong, Guangzhi, et al.
Published: (2026)
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
by: Hu, He, et al.
Published: (2026)
by: Hu, He, et al.
Published: (2026)
FindingEmo: An Image Dataset for Emotion Recognition in the Wild
by: Mertens, Laurent, et al.
Published: (2024)
by: Mertens, Laurent, et al.
Published: (2024)
MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models
by: Chen, Tianwei, et al.
Published: (2026)
by: Chen, Tianwei, et al.
Published: (2026)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
Understanding Figurative Meaning through Explainable Visual Entailment
by: Saakyan, Arkadiy, et al.
Published: (2024)
by: Saakyan, Arkadiy, et al.
Published: (2024)
How to Train Your Long-Context Visual Document Model
by: Veselka, Austin
Published: (2026)
by: Veselka, Austin
Published: (2026)
EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
by: Kawasaki, Haruka, et al.
Published: (2026)
by: Kawasaki, Haruka, et al.
Published: (2026)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
by: Tartaglini, Alexa R., et al.
Published: (2025)
by: Tartaglini, Alexa R., et al.
Published: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
by: Wang, Dianyi, et al.
Published: (2025)
by: Wang, Dianyi, et al.
Published: (2025)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
by: Wu, Chengyue, et al.
Published: (2024)
by: Wu, Chengyue, et al.
Published: (2024)
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models
by: Li, Mukai, et al.
Published: (2024)
by: Li, Mukai, et al.
Published: (2024)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
by: Wang, Qiuchen, et al.
Published: (2025)
by: Wang, Qiuchen, et al.
Published: (2025)
Efficient Adaptation For Remote Sensing Visual Grounding
by: Moughnieh, Hasan, et al.
Published: (2025)
by: Moughnieh, Hasan, et al.
Published: (2025)
EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition
by: Boudouri, Yassine El, et al.
Published: (2025)
by: Boudouri, Yassine El, et al.
Published: (2025)
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
by: Lee, Jewon, et al.
Published: (2025)
by: Lee, Jewon, et al.
Published: (2025)
Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy
by: Huang, Jiahao, et al.
Published: (2026)
by: Huang, Jiahao, et al.
Published: (2026)
Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages
by: Gain, Baban, et al.
Published: (2023)
by: Gain, Baban, et al.
Published: (2023)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
by: Zhou, Kaiwen, et al.
Published: (2023)
by: Zhou, Kaiwen, et al.
Published: (2023)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
by: Lian, Niu, et al.
Published: (2026)
by: Lian, Niu, et al.
Published: (2026)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
by: Fu, Xingyu, et al.
Published: (2023)
by: Fu, Xingyu, et al.
Published: (2023)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
by: Zhang, Haowei, et al.
Published: (2026)
by: Zhang, Haowei, et al.
Published: (2026)
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
by: Fu, Xingyu, et al.
Published: (2024)
by: Fu, Xingyu, et al.
Published: (2024)
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness
by: Sun, Yueru, et al.
Published: (2026)
by: Sun, Yueru, et al.
Published: (2026)
Emotion Recognition in Signers
by: Funakoshi, Kotaro, et al.
Published: (2025)
by: Funakoshi, Kotaro, et al.
Published: (2025)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
by: Zhang, Juntian, et al.
Published: (2025)
by: Zhang, Juntian, et al.
Published: (2025)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
by: Sharma, Aditya, et al.
Published: (2024)
by: Sharma, Aditya, et al.
Published: (2024)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025)
by: Lin, Bin, et al.
Published: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
by: Zhou, Yuqi, et al.
Published: (2025)
by: Zhou, Yuqi, et al.
Published: (2025)
Similar Items
-
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
by: Lu, Xingyuan, et al.
Published: (2025) -
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
by: Wu, Chengfei, et al.
Published: (2025) -
Internalized Reasoning for Long-Context Visual Document Understanding
by: Veselka, Austin
Published: (2026) -
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
by: Xie, Hongxia, et al.
Published: (2024) -
EmoCAM: Toward Understanding What Drives CNN-based Emotion Recognition
by: Doulfoukar, Youssef, et al.
Published: (2024)