Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhifeng, Zhang, Qixuan, Zhang, Peter, Niu, Wenjia, Zhang, Kaihao, Sankaranarayana, Ramesh, Caldwell, Sabrina, Gedeon, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Prompting in LLMs for Enhancing Emotion Recognition
by: Zhang, Qixuan, et al.
Published: (2024)
by: Zhang, Qixuan, et al.
Published: (2024)
LLDif: Diffusion Models for Low-light Emotion Recognition
by: Wang, Zhifeng, et al.
Published: (2024)
by: Wang, Zhifeng, et al.
Published: (2024)
LRDif: Diffusion Models for Under-Display Camera Emotion Recognition
by: Wang, Zhifeng, et al.
Published: (2024)
by: Wang, Zhifeng, et al.
Published: (2024)
Authentic Emotion Mapping: Benchmarking Facial Expressions in Real News
by: Zhang, Qixuan, et al.
Published: (2024)
by: Zhang, Qixuan, et al.
Published: (2024)
AstroRAG -- A Pagerank-Based Retrieval-Augmented Generation Pipeline for Question Answering in Astronomy
by: Wang, Zhifeng, et al.
Published: (2026)
by: Wang, Zhifeng, et al.
Published: (2026)
Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs
by: Yao, Yue, et al.
Published: (2025)
by: Yao, Yue, et al.
Published: (2025)
PromptRR: Diffusion Models as Prompt Generators for Single Image Reflection Removal
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs
by: Wang, Yahong, et al.
Published: (2025)
by: Wang, Yahong, et al.
Published: (2025)
TCP:Textual-based Class-aware Prompt tuning for Visual-Language Model
by: Yao, Hantao, et al.
Published: (2023)
by: Yao, Hantao, et al.
Published: (2023)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
by: Xenos, Alexandros, et al.
Published: (2024)
by: Xenos, Alexandros, et al.
Published: (2024)
Taylor Videos for Action Recognition
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
by: He, Wen-Jue, et al.
Published: (2025)
by: He, Wen-Jue, et al.
Published: (2025)
Visual Textualization for Image Prompted Object Detection
by: Wu, Yongjian, et al.
Published: (2025)
by: Wu, Yongjian, et al.
Published: (2025)
MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures
by: Morin, Lucas, et al.
Published: (2025)
by: Morin, Lucas, et al.
Published: (2025)
CAT: Coordinating Anatomical-Textual Prompts for Multi-Organ and Tumor Segmentation
by: Huang, Zhongzhen, et al.
Published: (2024)
by: Huang, Zhongzhen, et al.
Published: (2024)
Line of Sight: On Linear Representations in VLLMs
by: Rajaram, Achyuta, et al.
Published: (2025)
by: Rajaram, Achyuta, et al.
Published: (2025)
Synergistic Prompting for Robust Visual Recognition with Missing Modalities
by: Zhang, Zhihui, et al.
Published: (2025)
by: Zhang, Zhihui, et al.
Published: (2025)
Enhanced Textual Feature Extraction for Visual Question Answering: A Simple Convolutional Approach
by: Zhang, Zhilin, et al.
Published: (2024)
by: Zhang, Zhilin, et al.
Published: (2024)
Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild
by: Richet, Nicolas, et al.
Published: (2024)
by: Richet, Nicolas, et al.
Published: (2024)
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
by: QI, Anbin, et al.
Published: (2024)
by: QI, Anbin, et al.
Published: (2024)
When Spatial meets Temporal in Action Recognition
by: Chen, Huilin, et al.
Published: (2024)
by: Chen, Huilin, et al.
Published: (2024)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
by: Kumar, Puneet, et al.
Published: (2022)
by: Kumar, Puneet, et al.
Published: (2022)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
by: Feng, X., et al.
Published: (2024)
by: Feng, X., et al.
Published: (2024)
Gems: Group Emotion Profiling Through Multimodal Situational Understanding
by: Kataria, Anubhav, et al.
Published: (2025)
by: Kataria, Anubhav, et al.
Published: (2025)
Medal S: Spatio-Textual Prompt Model for Medical Segmentation
by: Shi, Pengcheng, et al.
Published: (2025)
by: Shi, Pengcheng, et al.
Published: (2025)
SEP: Self-Enhanced Prompt Tuning for Visual-Language Model
by: Yao, Hantao, et al.
Published: (2024)
by: Yao, Hantao, et al.
Published: (2024)
Bipartite Mode Matching for Vision Training Set Search from a Hierarchical Data Server
by: Yao, Yue, et al.
Published: (2026)
by: Yao, Yue, et al.
Published: (2026)
Toward a Holistic Evaluation of Robustness in CLIP Models
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
Seeing Beyond Redundancy: Task Complexity's Role in Vision Token Specialization in VLLMs
by: Hannan, Darryl, et al.
Published: (2026)
by: Hannan, Darryl, et al.
Published: (2026)
TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps
by: Raj, Arjun, et al.
Published: (2024)
by: Raj, Arjun, et al.
Published: (2024)
Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual Recognition
by: Li, Mengke, et al.
Published: (2024)
by: Li, Mengke, et al.
Published: (2024)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition
by: Liu, Ran, et al.
Published: (2025)
by: Liu, Ran, et al.
Published: (2025)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
by: Shi, Tong, et al.
Published: (2024)
by: Shi, Tong, et al.
Published: (2024)
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
by: Rosi, Gabriele, et al.
Published: (2025)
by: Rosi, Gabriele, et al.
Published: (2025)
Textualize Visual Prompt for Image Editing via Diffusion Bridge
by: Xu, Pengcheng, et al.
Published: (2025)
by: Xu, Pengcheng, et al.
Published: (2025)
Similar Items
-
Visual Prompting in LLMs for Enhancing Emotion Recognition
by: Zhang, Qixuan, et al.
Published: (2024) -
LLDif: Diffusion Models for Low-light Emotion Recognition
by: Wang, Zhifeng, et al.
Published: (2024) -
LRDif: Diffusion Models for Under-Display Camera Emotion Recognition
by: Wang, Zhifeng, et al.
Published: (2024) -
Authentic Emotion Mapping: Benchmarking Facial Expressions in Real News
by: Zhang, Qixuan, et al.
Published: (2024) -
AstroRAG -- A Pagerank-Based Retrieval-Augmented Generation Pipeline for Question Answering in Astronomy
by: Wang, Zhifeng, et al.
Published: (2026)