3D Audio-Visual Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sokolov, Artem, Bhosale, Swapnil, Zhu, Xiatian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
von: Mao, Yuxin, et al.
Veröffentlicht: (2023)
von: Mao, Yuxin, et al.
Veröffentlicht: (2023)
Audio-Visual Instance Segmentation
von: Guo, Ruohao, et al.
Veröffentlicht: (2023)
von: Guo, Ruohao, et al.
Veröffentlicht: (2023)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Benchmarking Cross-Domain Audio-Visual Deception Detection
von: Guo, Xiaobao, et al.
Veröffentlicht: (2024)
von: Guo, Xiaobao, et al.
Veröffentlicht: (2024)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
von: Xu, Le, et al.
Veröffentlicht: (2025)
von: Xu, Le, et al.
Veröffentlicht: (2025)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
von: Oorloff, Trevine, et al.
Veröffentlicht: (2024)
von: Oorloff, Trevine, et al.
Veröffentlicht: (2024)
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
von: Yu, Fei, et al.
Veröffentlicht: (2024)
von: Yu, Fei, et al.
Veröffentlicht: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
von: Guo, Yuxin, et al.
Veröffentlicht: (2024)
von: Guo, Yuxin, et al.
Veröffentlicht: (2024)
Aligned Better, Listen Better for Audio-Visual Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Classifying Shelf Life Quality of Pineapples by Combining Audio and Visual Features
von: Jiang, Yi-Lu, et al.
Veröffentlicht: (2025)
von: Jiang, Yi-Lu, et al.
Veröffentlicht: (2025)
Comparative Analysis of Modality Fusion Approaches for Audio-Visual Person Identification and Verification
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning
von: Yu, RunLin, et al.
Veröffentlicht: (2024)
von: Yu, RunLin, et al.
Veröffentlicht: (2024)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation
von: Li, Kai, et al.
Veröffentlicht: (2023)
von: Li, Kai, et al.
Veröffentlicht: (2023)
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
von: Zhao, Yusheng, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
von: Du, Jiarong, et al.
Veröffentlicht: (2025)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification
von: Xiao, Zhenyuan, et al.
Veröffentlicht: (2024)
von: Xiao, Zhenyuan, et al.
Veröffentlicht: (2024)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models
von: Tseng, Yuan, et al.
Veröffentlicht: (2023)
von: Tseng, Yuan, et al.
Veröffentlicht: (2023)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
von: Chen, Honglie, et al.
Veröffentlicht: (2024)
von: Chen, Honglie, et al.
Veröffentlicht: (2024)
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
Audio-Visual Segmentation via Unlabeled Frame Exploitation
von: Liu, Jinxiang, et al.
Veröffentlicht: (2024)
von: Liu, Jinxiang, et al.
Veröffentlicht: (2024)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
von: PN, Aravinda Reddy, et al.
Veröffentlicht: (2024)
von: PN, Aravinda Reddy, et al.
Veröffentlicht: (2024)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
von: Yang, Qi, et al.
Veröffentlicht: (2024)
von: Yang, Qi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Contrastive Conditional Latent Diffusion for Audio-visual Segmentation
von: Mao, Yuxin, et al.
Veröffentlicht: (2023) -
Audio-Visual Instance Segmentation
von: Guo, Ruohao, et al.
Veröffentlicht: (2023) -
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024) -
Benchmarking Cross-Domain Audio-Visual Deception Detection
von: Guo, Xiaobao, et al.
Veröffentlicht: (2024) -
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)