Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Dengming, You, Weitao, Li, Jingxiong, Lin, Weishen, Shi, Wenda, Zhao, Xue, Zuo, Heda, Wu, Junxian, Sun, Lingyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
di: Wu, Junxian, et al.
Pubblicazione: (2025)
di: Wu, Junxian, et al.
Pubblicazione: (2025)
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
di: Zhang, Dengming, et al.
Pubblicazione: (2024)
di: Zhang, Dengming, et al.
Pubblicazione: (2024)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
di: Zuo, Heda, et al.
Pubblicazione: (2025)
di: Zuo, Heda, et al.
Pubblicazione: (2025)
Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
di: Jia, Yanhao, et al.
Pubblicazione: (2025)
di: Jia, Yanhao, et al.
Pubblicazione: (2025)
FonTS: Text Rendering with Typography and Style Controls
di: Shi, Wenda, et al.
Pubblicazione: (2024)
di: Shi, Wenda, et al.
Pubblicazione: (2024)
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
di: Chen, Kai, et al.
Pubblicazione: (2024)
di: Chen, Kai, et al.
Pubblicazione: (2024)
WordCon: Word-level Typography Control in Scene Text Rendering
di: Shi, Wenda, et al.
Pubblicazione: (2025)
di: Shi, Wenda, et al.
Pubblicazione: (2025)
AnySurf: Any Surface Generation with Directed Edge
di: Shi, Wenda, et al.
Pubblicazione: (2026)
di: Shi, Wenda, et al.
Pubblicazione: (2026)
Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration
di: Teng, Zhuyu, et al.
Pubblicazione: (2026)
di: Teng, Zhuyu, et al.
Pubblicazione: (2026)
Efficient and Scalable Chinese Vector Font Generation via Component Composition
di: Song, Jinyu, et al.
Pubblicazione: (2024)
di: Song, Jinyu, et al.
Pubblicazione: (2024)
See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025)
di: Nguyen, Le Thien Phuc, et al.
Pubblicazione: (2025)
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
di: Limpijankit, Marvin, et al.
Pubblicazione: (2026)
di: Limpijankit, Marvin, et al.
Pubblicazione: (2026)
EmoArt: A Multidimensional Dataset for Emotion-Aware Artistic Generation
di: Zhang, Cheng, et al.
Pubblicazione: (2025)
di: Zhang, Cheng, et al.
Pubblicazione: (2025)
The Audio-Visual BatVision Dataset for Research on Sight and Sound
di: Brunetto, Amandine, et al.
Pubblicazione: (2023)
di: Brunetto, Amandine, et al.
Pubblicazione: (2023)
Green Energy and State Power: The Case of Zhanatas Wind Power Project in Kazakhstan
di: Weishen Zeng
Pubblicazione: (2025)
di: Weishen Zeng
Pubblicazione: (2025)
It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
di: Park, Sooyoung, et al.
Pubblicazione: (2025)
di: Park, Sooyoung, et al.
Pubblicazione: (2025)
Multilevel constructions of constant dimension codes based on one-factorization of complete graphs
di: Xu, Dengming, et al.
Pubblicazione: (2025)
di: Xu, Dengming, et al.
Pubblicazione: (2025)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
di: Wang, Qiaolin, et al.
Pubblicazione: (2025)
di: Wang, Qiaolin, et al.
Pubblicazione: (2025)
Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
di: Wang, Haozhe, et al.
Pubblicazione: (2026)
Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition
di: Deng, Shijian, et al.
Pubblicazione: (2024)
di: Deng, Shijian, et al.
Pubblicazione: (2024)
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models
di: Lyu, Zesen, et al.
Pubblicazione: (2025)
di: Lyu, Zesen, et al.
Pubblicazione: (2025)
Viscometric investigations and molecular interactions of some derivatives of 5-substituted indole dihydropyrimidines in mixed organic solvents
di: L. C. Heda
Pubblicazione: (2010)
di: L. C. Heda
Pubblicazione: (2010)
Large Language Models Implicitly Learn to See and Hear Just By Reading
di: Verma, Prateek, et al.
Pubblicazione: (2025)
di: Verma, Prateek, et al.
Pubblicazione: (2025)
Do Audio-Visual Large Language Models Really See and Hear?
di: Selvakumar, Ramaneswaran, et al.
Pubblicazione: (2026)
di: Selvakumar, Ramaneswaran, et al.
Pubblicazione: (2026)
HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation
di: Yang, Hongji, et al.
Pubblicazione: (2026)
di: Yang, Hongji, et al.
Pubblicazione: (2026)
Vision Language Models See What You Want but not What You See
di: Gao, Qingying, et al.
Pubblicazione: (2024)
di: Gao, Qingying, et al.
Pubblicazione: (2024)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
di: Sun, Boyuan, et al.
Pubblicazione: (2026)
di: Sun, Boyuan, et al.
Pubblicazione: (2026)
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
di: Yuan, Zhengqing, et al.
Pubblicazione: (2023)
di: Yuan, Zhengqing, et al.
Pubblicazione: (2023)
See Me, Hear Me: Skype in the Classroom
di: Foote, Carolyn
Pubblicazione: (2008)
di: Foote, Carolyn
Pubblicazione: (2008)
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
di: Shi, Zhengpeng, et al.
Pubblicazione: (2025)
di: Shi, Zhengpeng, et al.
Pubblicazione: (2025)
Spray Coating of Thick Perovskite Films for Photodetectors: The Aerosol–Liquid–Solid Mechanisms and Sensing Applications
di: Wei Qian, et al.
Pubblicazione: (2026)
di: Wei Qian, et al.
Pubblicazione: (2026)
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
di: Srikrishnan, Tharun Adithya, et al.
Pubblicazione: (2025)
di: Srikrishnan, Tharun Adithya, et al.
Pubblicazione: (2025)
InsightSee: Advancing Multi-agent Vision-Language Models for Enhanced Visual Understanding
di: Zhang, Huaxiang, et al.
Pubblicazione: (2024)
di: Zhang, Huaxiang, et al.
Pubblicazione: (2024)
Meta-aware Learning in text-to-SQL Large Language Model
di: Zhang, Wenda
Pubblicazione: (2025)
di: Zhang, Wenda
Pubblicazione: (2025)
UniHetero: Could Generation Enhance Understanding for Vision-Language-Model at Large Data Scale?
di: Chen, Fengjiao, et al.
Pubblicazione: (2025)
di: Chen, Fengjiao, et al.
Pubblicazione: (2025)
EgoSound: Benchmarking Sound Understanding in Egocentric Videos
di: Zhu, Bingwen, et al.
Pubblicazione: (2026)
di: Zhu, Bingwen, et al.
Pubblicazione: (2026)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
di: Zhang, Jialiang, et al.
Pubblicazione: (2026)
di: Zhang, Jialiang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
di: Wu, Junxian, et al.
Pubblicazione: (2025) -
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
di: Zhang, Dengming, et al.
Pubblicazione: (2024) -
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
di: Zuo, Heda, et al.
Pubblicazione: (2025) -
Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization
di: Jia, Yanhao, et al.
Pubblicazione: (2025) -
FonTS: Text Rendering with Typography and Style Controls
di: Shi, Wenda, et al.
Pubblicazione: (2024)