FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Weidong, Ye, Cheng, Mao, Zhendong, Song, Peipei, Liu, Xinyan, Zhang, Lei, Chang, Xiaojun, Zhang, Yongdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dual-path Collaborative Generation Network for Emotional Video Captioning
von: Ye, Cheng, et al.
Veröffentlicht: (2024)
von: Ye, Cheng, et al.
Veröffentlicht: (2024)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
von: Guo, Yijie, et al.
Veröffentlicht: (2025)
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
von: Chen, Weidong, et al.
Veröffentlicht: (2026)
von: Chen, Weidong, et al.
Veröffentlicht: (2026)
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation
von: Wang, Liping, et al.
Veröffentlicht: (2026)
von: Wang, Liping, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Egocentric Video Captioning
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering
von: Song, Peipei, et al.
Veröffentlicht: (2025)
von: Song, Peipei, et al.
Veröffentlicht: (2025)
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
von: Xia, Hou, et al.
Veröffentlicht: (2025)
von: Xia, Hou, et al.
Veröffentlicht: (2025)
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
von: Hu, Jinpeng, et al.
Veröffentlicht: (2025)
von: Hu, Jinpeng, et al.
Veröffentlicht: (2025)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
von: Wu, Bin, et al.
Veröffentlicht: (2026)
von: Wu, Bin, et al.
Veröffentlicht: (2026)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
von: Huang, Mengqi, et al.
Veröffentlicht: (2024)
von: Huang, Mengqi, et al.
Veröffentlicht: (2024)
DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
von: Huang, Mengqi, et al.
Veröffentlicht: (2022)
von: Huang, Mengqi, et al.
Veröffentlicht: (2022)
Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting
von: Fu, Fengyi, et al.
Veröffentlicht: (2024)
von: Fu, Fengyi, et al.
Veröffentlicht: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
von: Jeon, MinJu, et al.
Veröffentlicht: (2025)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
von: Zhang, Long, et al.
Veröffentlicht: (2025)
von: Zhang, Long, et al.
Veröffentlicht: (2025)
SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities
von: Faghihi, Ehsan, et al.
Veröffentlicht: (2024)
von: Faghihi, Ehsan, et al.
Veröffentlicht: (2024)
Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
von: Han, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Han, Zhiyuan, et al.
Veröffentlicht: (2025)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
von: Rosa, Kevin Dela
Veröffentlicht: (2024)
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
von: Wang, Zeheng, et al.
Veröffentlicht: (2026)
von: Wang, Zeheng, et al.
Veröffentlicht: (2026)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
von: Mao, Zhendong, et al.
Veröffentlicht: (2024)
von: Mao, Zhendong, et al.
Veröffentlicht: (2024)
Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights
von: Mao, Shunqi, et al.
Veröffentlicht: (2024)
von: Mao, Shunqi, et al.
Veröffentlicht: (2024)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
von: Li, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2023)
Composed Multi-modal Retrieval: A Survey of Approaches and Applications
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2025)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
FACE: Few-shot Adapter with Cross-view Fusion for Cross-subject EEG Emotion Recognition
von: Liu, Haiqi, et al.
Veröffentlicht: (2025)
von: Liu, Haiqi, et al.
Veröffentlicht: (2025)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
Open Multimodal Retrieval-Augmented Factual Image Generation
von: Tian, Yang, et al.
Veröffentlicht: (2025)
von: Tian, Yang, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
von: Cioni, Dario, et al.
Veröffentlicht: (2023)
von: Cioni, Dario, et al.
Veröffentlicht: (2023)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
von: Long, Xiaosheng, et al.
Veröffentlicht: (2025)
von: Long, Xiaosheng, et al.
Veröffentlicht: (2025)
Taming Transformer for Emotion-Controllable Talking Face Generation
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziqi, et al.
Veröffentlicht: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities
von: Fan, Qi, et al.
Veröffentlicht: (2024)
von: Fan, Qi, et al.
Veröffentlicht: (2024)
Self-supervised Gait-based Emotion Representation Learning from Selective Strongly Augmented Skeleton Sequences
von: Song, Cheng, et al.
Veröffentlicht: (2024)
von: Song, Cheng, et al.
Veröffentlicht: (2024)
Federated Dialogue-Semantic Diffusion for Emotion Recognition under Incomplete Modalities
von: Qiu, Xihang, et al.
Veröffentlicht: (2025)
von: Qiu, Xihang, et al.
Veröffentlicht: (2025)
Beyond Caption-Based Queries for Video Moment Retrieval
von: Pujol-Perich, David, et al.
Veröffentlicht: (2026)
von: Pujol-Perich, David, et al.
Veröffentlicht: (2026)
Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions
von: Sun, Licai, et al.
Veröffentlicht: (2025)
von: Sun, Licai, et al.
Veröffentlicht: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
von: Yao, Linli, et al.
Veröffentlicht: (2026)
von: Yao, Linli, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dual-path Collaborative Generation Network for Emotional Video Captioning
von: Ye, Cheng, et al.
Veröffentlicht: (2024) -
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
von: Guo, Yijie, et al.
Veröffentlicht: (2025) -
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
von: Chen, Weidong, et al.
Veröffentlicht: (2026) -
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation
von: Wang, Liping, et al.
Veröffentlicht: (2026) -
Retrieval-Augmented Egocentric Video Captioning
von: Xu, Jilan, et al.
Veröffentlicht: (2024)