OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Shengkai, Yin, Yifang, Cao, Jinming, Xiang, Shili, Liu, Zhenguang, Zimmermann, Roger |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open-Vocabulary Audio-Visual Semantic Segmentation
von: Guo, Ruohao, et al.
Veröffentlicht: (2024)
von: Guo, Ruohao, et al.
Veröffentlicht: (2024)
Towards Open-Vocabulary Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Towards Open-Vocabulary Video Semantic Segmentation
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
von: Ye, Chengyang, et al.
Veröffentlicht: (2024)
von: Ye, Chengyang, et al.
Veröffentlicht: (2024)
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
AcoustEmo: Open-Vocabulary Emotion Reasoning via Utterance-Aware Acoustic Q-Former
von: Zhang, Liyun, et al.
Veröffentlicht: (2026)
von: Zhang, Liyun, et al.
Veröffentlicht: (2026)
Audio-Visual Instance Segmentation
von: Guo, Ruohao, et al.
Veröffentlicht: (2023)
von: Guo, Ruohao, et al.
Veröffentlicht: (2023)
Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits
von: Ishii, Masato, et al.
Veröffentlicht: (2025)
von: Ishii, Masato, et al.
Veröffentlicht: (2025)
OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
von: Fang, Pengcheng, et al.
Veröffentlicht: (2026)
von: Fang, Pengcheng, et al.
Veröffentlicht: (2026)
Network Bending of Diffusion Models for Audio-Visual Generation
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
von: Wang, Yaoting, et al.
Veröffentlicht: (2023)
von: Wang, Yaoting, et al.
Veröffentlicht: (2023)
PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models
von: Wnag, Zining, et al.
Veröffentlicht: (2024)
von: Wnag, Zining, et al.
Veröffentlicht: (2024)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2025)
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2025)
Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
von: Zhou, Jinxing, et al.
Veröffentlicht: (2026)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2026)
Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models
von: Wen, Jinming, et al.
Veröffentlicht: (2025)
von: Wen, Jinming, et al.
Veröffentlicht: (2025)
EEG2TEXT-CN: An Exploratory Study of Open-Vocabulary Chinese Text-EEG Alignment via Large Language Model and Contrastive Learning on ChineseEEG
von: Lu, Jacky Tai-Yu, et al.
Veröffentlicht: (2025)
von: Lu, Jacky Tai-Yu, et al.
Veröffentlicht: (2025)
Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation
von: Chen, Yuanhong, et al.
Veröffentlicht: (2023)
von: Chen, Yuanhong, et al.
Veröffentlicht: (2023)
Recent Advances of End-to-End Video Coding Technologies for AVS Standard Development
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
von: Zhu, Sa, et al.
Veröffentlicht: (2026)
von: Zhu, Sa, et al.
Veröffentlicht: (2026)
Collaborative Hybrid Propagator for Temporal Misalignment in Audio-Visual Segmentation
von: Li, Kexin, et al.
Veröffentlicht: (2024)
von: Li, Kexin, et al.
Veröffentlicht: (2024)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
von: Wu, Xiongwei, et al.
Veröffentlicht: (2024)
Multimodal Information Retrieval for Open World with Edit Distance Weak Supervision
von: Solaiman, KMA, et al.
Veröffentlicht: (2025)
von: Solaiman, KMA, et al.
Veröffentlicht: (2025)
MCAD: Multimodal Context-Aware Audio Description Generation For Soccer
von: Chaudhary, Lipisha, et al.
Veröffentlicht: (2025)
von: Chaudhary, Lipisha, et al.
Veröffentlicht: (2025)
MCIGLE: Multimodal Exemplar-Free Class-Incremental Graph Learning
von: You, Haochen, et al.
Veröffentlicht: (2025)
von: You, Haochen, et al.
Veröffentlicht: (2025)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation
von: Su, Kun, et al.
Veröffentlicht: (2024)
von: Su, Kun, et al.
Veröffentlicht: (2024)
Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings
von: Liu, Shimiao, et al.
Veröffentlicht: (2025)
von: Liu, Shimiao, et al.
Veröffentlicht: (2025)
PetalView: Fine-grained Location and Orientation Extraction of Street-view Images via Cross-view Local Search with Supplementary Materials
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
von: Hu, Wenmiao, et al.
Veröffentlicht: (2024)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
Look, Listen and Segment: Towards Weakly Supervised Audio-visual Semantic Segmentation
von: Li, Chengzhi, et al.
Veröffentlicht: (2026)
von: Li, Chengzhi, et al.
Veröffentlicht: (2026)
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor
von: Sun, Shuyang, et al.
Veröffentlicht: (2023)
von: Sun, Shuyang, et al.
Veröffentlicht: (2023)
IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models
von: Rahimi, Hamed, et al.
Veröffentlicht: (2026)
von: Rahimi, Hamed, et al.
Veröffentlicht: (2026)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
von: Zou, Heqing, et al.
Veröffentlicht: (2024)
von: Zou, Heqing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Open-Vocabulary Audio-Visual Semantic Segmentation
von: Guo, Ruohao, et al.
Veröffentlicht: (2024) -
Towards Open-Vocabulary Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024) -
Towards Open-Vocabulary Video Semantic Segmentation
von: Li, Xinhao, et al.
Veröffentlicht: (2024) -
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
von: Ye, Chengyang, et al.
Veröffentlicht: (2024) -
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
von: Li, Jia, et al.
Veröffentlicht: (2025)