Gespeichert in:
| Hauptverfasser: | Zhao, Jingwei, Wen, Yuhua, Li, Qifei, Hu, Minchi, Zhou, Yingying, Xue, Jingyao, Wu, Junyang, Gao, Yingming, Wen, Zhengqi, Tao, Jianhua, Li, Ya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.22934 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
von: Wen, Yuhua, et al.
Veröffentlicht: (2025)
von: Wen, Yuhua, et al.
Veröffentlicht: (2025)
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
von: Li, Qifei, et al.
Veröffentlicht: (2024)
von: Li, Qifei, et al.
Veröffentlicht: (2024)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
Spatio-Temporal Cluster-Triggered Encoding for Spiking Neural Networks
von: Hu, Minchi
Veröffentlicht: (2025)
von: Hu, Minchi
Veröffentlicht: (2025)
Psy-Insight: Explainable Multi-turn Bilingual Dataset for Mental Health Counseling
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
Exploring the Role of Audio in Multimodal Misinformation Detection
von: Liu, Moyang, et al.
Veröffentlicht: (2024)
von: Liu, Moyang, et al.
Veröffentlicht: (2024)
PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis
von: Xie, Heng, et al.
Veröffentlicht: (2025)
von: Xie, Heng, et al.
Veröffentlicht: (2025)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
COAL: Robust Contrastive Learning‐Based Visual Navigation Framework
von: Zengmao Wang, et al.
Veröffentlicht: (2025)
von: Zengmao Wang, et al.
Veröffentlicht: (2025)
Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning
von: Yan, Kaiying, et al.
Veröffentlicht: (2025)
von: Yan, Kaiying, et al.
Veröffentlicht: (2025)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement
von: Yang, Zhengxian, et al.
Veröffentlicht: (2026)
von: Yang, Zhengxian, et al.
Veröffentlicht: (2026)
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
von: Zhang, Haojie, et al.
Veröffentlicht: (2024)
von: Zhang, Haojie, et al.
Veröffentlicht: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
Fake News Detection and Manipulation Reasoning via Large Vision-Language Models
von: Jin, Ruihan, et al.
Veröffentlicht: (2024)
von: Jin, Ruihan, et al.
Veröffentlicht: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
Residual Speaker Representation for One-Shot Voice Conversion
von: Xu, Le, et al.
Veröffentlicht: (2023)
von: Xu, Le, et al.
Veröffentlicht: (2023)
Deep Smart Contract Intent Detection
von: Huang, Youwei, et al.
Veröffentlicht: (2022)
von: Huang, Youwei, et al.
Veröffentlicht: (2022)
A Survey on Backbones for Deep Video Action Recognition
von: Tang, Zixuan, et al.
Veröffentlicht: (2024)
von: Tang, Zixuan, et al.
Veröffentlicht: (2024)
Explainable Multimodal Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
von: Lian, Zheng, et al.
Veröffentlicht: (2023)
ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
von: Yan, Kaiying, et al.
Veröffentlicht: (2025)
von: Yan, Kaiying, et al.
Veröffentlicht: (2025)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2026)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2026)
RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing
von: Jin, Ruihan, et al.
Veröffentlicht: (2025)
von: Jin, Ruihan, et al.
Veröffentlicht: (2025)
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
von: Wu, Jinyang, et al.
Veröffentlicht: (2026)
ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
von: Yang, Zhengxian, et al.
Veröffentlicht: (2025)
von: Yang, Zhengxian, et al.
Veröffentlicht: (2025)
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty
von: Zhang, Zhen, et al.
Veröffentlicht: (2024)
von: Zhang, Zhen, et al.
Veröffentlicht: (2024)
TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression Recognition
von: Zhu, Jianhua, et al.
Veröffentlicht: (2024)
von: Zhu, Jianhua, et al.
Veröffentlicht: (2024)
Generalized Fake Audio Detection via Deep Stable Learning
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
von: Lian, Zheng, et al.
Veröffentlicht: (2024)
von: Lian, Zheng, et al.
Veröffentlicht: (2024)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
EvoSpark: Endogenous Interactive Agent Societies for Unified Long-Horizon Narrative Evolution
von: He, Shiyu, et al.
Veröffentlicht: (2026)
von: He, Shiyu, et al.
Veröffentlicht: (2026)
Psy-Copilot: Visual Chain of Thought for Counseling
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
von: Chen, Keqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DashFusion: Dual-stream Alignment with Hierarchical Bottleneck Fusion for Multimodal Sentiment Analysis
von: Wen, Yuhua, et al.
Veröffentlicht: (2025) -
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
von: Li, Qifei, et al.
Veröffentlicht: (2024) -
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
von: Wang, Cong, et al.
Veröffentlicht: (2025) -
Spatio-Temporal Cluster-Triggered Encoding for Spiking Neural Networks
von: Hu, Minchi
Veröffentlicht: (2025) -
Psy-Insight: Explainable Multi-turn Bilingual Dataset for Mental Health Counseling
von: Chen, Keqi, et al.
Veröffentlicht: (2025)