Gespeichert in:
| Hauptverfasser: | Zhao, Jinghua, Jia, Yuhang, Wang, Shiyao, Zhou, Jiaming, Wang, Hui, Qin, Yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2504.15066 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Versatile audio-visual learning for emotion recognition
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
VCEMO: Multi-Modal Emotion Recognition for Chinese Voiceprints
von: Tang, Jinghua, et al.
Veröffentlicht: (2024)
von: Tang, Jinghua, et al.
Veröffentlicht: (2024)
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
von: Xu, Pengju, et al.
Veröffentlicht: (2025)
von: Xu, Pengju, et al.
Veröffentlicht: (2025)
ConCLVD: Controllable Chinese Landscape Video Generation via Diffusion Model
von: Liu, Dingming, et al.
Veröffentlicht: (2024)
von: Liu, Dingming, et al.
Veröffentlicht: (2024)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
PoEmotion: Can AI Utilize Chinese Calligraphy to Express Emotion from Poems?
von: Liu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Liu, Tiancheng, et al.
Veröffentlicht: (2025)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
von: Li, Cancan, et al.
Veröffentlicht: (2025)
von: Li, Cancan, et al.
Veröffentlicht: (2025)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Archiving Body Movements: Collective Generation of Chinese Calligraphy
von: Zhou, Aven Le, et al.
Veröffentlicht: (2023)
von: Zhou, Aven Le, et al.
Veröffentlicht: (2023)
A Multimodal Transformer for Live Streaming Highlight Prediction
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
von: Deng, Jiaxin, et al.
Veröffentlicht: (2024)
Chinese for beginners [conjunto] = Ru men han yu / Sun Jun
von: Jun Sun
von: Jun Sun
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
Semi-supervised Chinese Poem-to-Painting Generation via Cycle-consistent Adversarial Networks
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
Visual-based spatial audio generation system for multi-speaker environments
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
GestureHYDRA: Semantic Co-speech Gesture Synthesis via Hybrid Modality Diffusion Transformer and Cascaded-Synchronized Retrieval-Augmented Generation
von: Yang, Quanwei, et al.
Veröffentlicht: (2025)
von: Yang, Quanwei, et al.
Veröffentlicht: (2025)
Shorter Is Different: Characterizing the Dynamics of Short-Form Video Platforms
von: Chen, Zhilong, et al.
Veröffentlicht: (2024)
von: Chen, Zhilong, et al.
Veröffentlicht: (2024)
CPCLDETECTOR: Knowledge Enhancement and Alignment Selection for Chinese Patronizing and Condescending Language Detection
von: Yang, Jiaxun, et al.
Veröffentlicht: (2025)
von: Yang, Jiaxun, et al.
Veröffentlicht: (2025)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
von: Yin, Yongkang, et al.
Veröffentlicht: (2023)
von: Yin, Yongkang, et al.
Veröffentlicht: (2023)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation
von: Ling, Zeyu, et al.
Veröffentlicht: (2025)
von: Ling, Zeyu, et al.
Veröffentlicht: (2025)
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
von: Wang, Yingna, et al.
Veröffentlicht: (2025)
von: Wang, Yingna, et al.
Veröffentlicht: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
Look, Listen and Segment: Towards Weakly Supervised Audio-visual Semantic Segmentation
von: Li, Chengzhi, et al.
Veröffentlicht: (2026)
von: Li, Chengzhi, et al.
Veröffentlicht: (2026)
Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
von: Liu, Huijie, et al.
Veröffentlicht: (2025)
Chinese idioms [juego] = Cheng yu gu shi
Towards Automatic Soccer Commentary Generation with Knowledge-Enhanced Visual Reasoning
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
von: Jin, Zeyu, et al.
Veröffentlicht: (2026)
Fully Automatic Content-Aware Tiling Pipeline for Pathology Whole Slide Images
von: Jabar, Falah, et al.
Veröffentlicht: (2024)
von: Jabar, Falah, et al.
Veröffentlicht: (2024)
Resolution deficits drive simulator sickness and compromise reading performance in virtual environments
von: Wang, Jialin, et al.
Veröffentlicht: (2026)
von: Wang, Jialin, et al.
Veröffentlicht: (2026)
DiffBrush:Just Painting the Art by Your Hands
von: Chu, Jiaming, et al.
Veröffentlicht: (2025)
von: Chu, Jiaming, et al.
Veröffentlicht: (2025)
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
von: Qiu, Ke, et al.
Veröffentlicht: (2026)
von: Qiu, Ke, et al.
Veröffentlicht: (2026)
Task Presentation and Human Perception in Interactive Video Retrieval
von: Willis, Nina, et al.
Veröffentlicht: (2024)
von: Willis, Nina, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
von: Sun, Haoqin, et al.
Veröffentlicht: (2025) -
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026) -
Versatile audio-visual learning for emotion recognition
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023) -
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024) -
VCEMO: Multi-Modal Emotion Recognition for Chinese Voiceprints
von: Tang, Jinghua, et al.
Veröffentlicht: (2024)