Gespeichert in:
| Hauptverfasser: | Cong, Kaixuan, Wang, Yifan, Xue, Rongkun, Jiang, Yuyang, Feng, Yiming, Yang, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.09323 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cs2K: Class-specific and Class-shared Knowledge Guidance for Incremental Semantic Segmentation
von: Cong, Wei, et al.
Veröffentlicht: (2024)
von: Cong, Wei, et al.
Veröffentlicht: (2024)
Inconsistency-Aware Cross-Attention for Audio-Visual Fusion in Dimensional Emotion Recognition
von: Rajasekhar, G, et al.
Veröffentlicht: (2024)
von: Rajasekhar, G, et al.
Veröffentlicht: (2024)
RFPPO: Motion Dynamic RRT based Fluid Field - PPO for Dynamic TF/TA Routing Planning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
von: Shi, Tong, et al.
Veröffentlicht: (2024)
von: Shi, Tong, et al.
Veröffentlicht: (2024)
FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
von: Li, Wei, et al.
Veröffentlicht: (2026)
von: Li, Wei, et al.
Veröffentlicht: (2026)
Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition
von: Zuo, Yukun, et al.
Veröffentlicht: (2024)
von: Zuo, Yukun, et al.
Veröffentlicht: (2024)
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
von: Ye, Qilang, et al.
Veröffentlicht: (2025)
von: Ye, Qilang, et al.
Veröffentlicht: (2025)
Contrastive Learning for Multimodal Human Activity Recognition with Limited Labeled Data
von: Jing, Long, et al.
Veröffentlicht: (2026)
von: Jing, Long, et al.
Veröffentlicht: (2026)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
von: Xue, Rongkun, et al.
Veröffentlicht: (2024)
InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
ConfusionBench: An Expert-Validated Benchmark for Confusion Recognition and Localization in Educational Videos
von: Dong, Lu, et al.
Veröffentlicht: (2026)
von: Dong, Lu, et al.
Veröffentlicht: (2026)
Context-Aware Aerial Object Detection: Leveraging Inter-Object and Background Relationships
von: Ren, Botao, et al.
Veröffentlicht: (2024)
von: Ren, Botao, et al.
Veröffentlicht: (2024)
Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
von: Abid, Abderrazek, et al.
Veröffentlicht: (2025)
von: Abid, Abderrazek, et al.
Veröffentlicht: (2025)
eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
von: Wu, Xuecheng, et al.
Veröffentlicht: (2025)
von: Wu, Xuecheng, et al.
Veröffentlicht: (2025)
Dynamic Attention and Bi-directional Fusion for Safety Helmet Wearing Detection
von: Feng, Junwei, et al.
Veröffentlicht: (2024)
von: Feng, Junwei, et al.
Veröffentlicht: (2024)
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
von: Hooshanfar, Kiana, et al.
Veröffentlicht: (2025)
von: Hooshanfar, Kiana, et al.
Veröffentlicht: (2025)
Understanding Open-Set Recognition by Jacobian Norm and Inter-Class Separation
von: Park, Jaewoo, et al.
Veröffentlicht: (2022)
von: Park, Jaewoo, et al.
Veröffentlicht: (2022)
Noise-Tolerant Learning for Audio-Visual Action Recognition
von: Han, Haochen, et al.
Veröffentlicht: (2022)
von: Han, Haochen, et al.
Veröffentlicht: (2022)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
von: Yu, Yinfeng, et al.
Veröffentlicht: (2025)
Evaluating Attribute Confusion in Fashion Text-to-Image Generation
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
von: Liu, Ziyue, et al.
Veröffentlicht: (2025)
Cross Attentional Audio-Visual Fusion for Dimensional Emotion Recognition
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2021)
von: Praveen, R. Gnana, et al.
Veröffentlicht: (2021)
The Comparability of Model Fusion to Measured Data in Confuser Rejection
von: Flynn, Conor, et al.
Veröffentlicht: (2025)
von: Flynn, Conor, et al.
Veröffentlicht: (2025)
DyCAF-Net: Dynamic Class-Aware Fusion Network
von: Jahin, Md Abrar, et al.
Veröffentlicht: (2025)
von: Jahin, Md Abrar, et al.
Veröffentlicht: (2025)
Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
von: She, Yifei, et al.
Veröffentlicht: (2025)
von: She, Yifei, et al.
Veröffentlicht: (2025)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
von: Liang, Sen, et al.
Veröffentlicht: (2026)
von: Liang, Sen, et al.
Veröffentlicht: (2026)
Confidence-driven Gradient Modulation for Multimodal Human Activity Recognition: A Dynamic Contrastive Dual-Path Learning Approach
von: Ji, Panpan, et al.
Veröffentlicht: (2025)
von: Ji, Panpan, et al.
Veröffentlicht: (2025)
Efficient Audio-Visual Fusion for Video Classification
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
Universal Incremental Learning: Mitigating Confusion from Inter- and Intra-task Distribution Randomness
von: Luo, Sheng, et al.
Veröffentlicht: (2025)
von: Luo, Sheng, et al.
Veröffentlicht: (2025)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
von: Li, Huilai, et al.
Veröffentlicht: (2025)
von: Li, Huilai, et al.
Veröffentlicht: (2025)
CLASP: Class-Adaptive Layer Fusion and Dual-Stage Pruning for Multimodal Large Language Models
von: Dang, Yunkai, et al.
Veröffentlicht: (2026)
von: Dang, Yunkai, et al.
Veröffentlicht: (2026)
WiFi based Human Fall and Activity Recognition using Transformer based Encoder Decoder and Graph Neural Networks
von: Cho, Younggeol, et al.
Veröffentlicht: (2025)
von: Cho, Younggeol, et al.
Veröffentlicht: (2025)
Triple Spectral Fusion for Sensor-based Human Activity Recognition
von: Zhang, Ye, et al.
Veröffentlicht: (2026)
von: Zhang, Ye, et al.
Veröffentlicht: (2026)
AdaFedFR: Federated Face Recognition with Adaptive Inter-Class Representation Learning
von: Qiu, Di, et al.
Veröffentlicht: (2024)
von: Qiu, Di, et al.
Veröffentlicht: (2024)
Class-Aware Mask-Guided Feature Refinement for Scene Text Recognition
von: Yang, Mingkun, et al.
Veröffentlicht: (2024)
von: Yang, Mingkun, et al.
Veröffentlicht: (2024)
IDSelect: A RL-Based Cost-Aware Selection Agent for Video-based Multi-Modal Person Recognition
von: Ji, Yuyang, et al.
Veröffentlicht: (2026)
von: Ji, Yuyang, et al.
Veröffentlicht: (2026)
InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
von: Wu, Zizhao, et al.
Veröffentlicht: (2025)
von: Wu, Zizhao, et al.
Veröffentlicht: (2025)
Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
von: Awan, Mahrukh, et al.
Veröffentlicht: (2024)
Visual Confused Deputy: Exploiting and Defending Perception Failures in Computer-Using Agents
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities
von: Song, Peibo, et al.
Veröffentlicht: (2026)
von: Song, Peibo, et al.
Veröffentlicht: (2026)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cs2K: Class-specific and Class-shared Knowledge Guidance for Incremental Semantic Segmentation
von: Cong, Wei, et al.
Veröffentlicht: (2024) -
Inconsistency-Aware Cross-Attention for Audio-Visual Fusion in Dimensional Emotion Recognition
von: Rajasekhar, G, et al.
Veröffentlicht: (2024) -
RFPPO: Motion Dynamic RRT based Fluid Field - PPO for Dynamic TF/TA Routing Planning
von: Xue, Rongkun, et al.
Veröffentlicht: (2024) -
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
von: Shi, Tong, et al.
Veröffentlicht: (2024) -
FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
von: Li, Wei, et al.
Veröffentlicht: (2026)