Unsupervised Audio-Visual Segmentation with Modality Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Bhosale, Swapnil, Yang, Haosen, Kanojia, Diptesh, Deng, Jiangkang, Zhu, Xiatian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
by: Bhosale, Swapnil, et al.
Published: (2024)
by: Bhosale, Swapnil, et al.
Published: (2024)
3D Audio-Visual Segmentation
by: Sokolov, Artem, et al.
Published: (2024)
by: Sokolov, Artem, et al.
Published: (2024)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024)
by: Nazarieh, Fatemeh, et al.
Published: (2024)
Recognize Any Regions
by: Yang, Haosen, et al.
Published: (2023)
by: Yang, Haosen, et al.
Published: (2023)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
Rethinking Alignment and Uniformity in Unsupervised Semantic Segmentation
by: Zhang, Daoan, et al.
Published: (2022)
by: Zhang, Daoan, et al.
Published: (2022)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
by: Zhou, Ziwei, et al.
Published: (2025)
by: Zhou, Ziwei, et al.
Published: (2025)
U-Mamba2-SSL for Semi-Supervised Tooth and Pulp Segmentation in CBCT
by: Tan, Zhi Qin, et al.
Published: (2025)
by: Tan, Zhi Qin, et al.
Published: (2025)
U-Mamba2: Scaling State Space Models for Dental Anatomy Segmentation in CBCT
by: Tan, Zhi Qin, et al.
Published: (2025)
by: Tan, Zhi Qin, et al.
Published: (2025)
Uncertainty-Aware Pseudo-Label Filtering for Source-Free Unsupervised Domain Adaptation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Unsupervised Domain Adaptation via Similarity-based Prototypes for Cross-Modality Segmentation
by: Ye, Ziyu, et al.
Published: (2025)
by: Ye, Ziyu, et al.
Published: (2025)
TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
by: Liu, Xinran, et al.
Published: (2026)
by: Liu, Xinran, et al.
Published: (2026)
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
by: Nazarieh, Fatemeh, et al.
Published: (2025)
by: Nazarieh, Fatemeh, et al.
Published: (2025)
Few-Shot Medical Image Segmentation with High-Fidelity Prototypes
by: Tang, Song, et al.
Published: (2024)
by: Tang, Song, et al.
Published: (2024)
Audio Visual Segmentation Through Text Embeddings
by: Lee, Kyungbok, et al.
Published: (2025)
by: Lee, Kyungbok, et al.
Published: (2025)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided Alignment
by: Chen, Xintao, et al.
Published: (2025)
by: Chen, Xintao, et al.
Published: (2025)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
MIAR: Modality Interaction and Alignment Representation Fuison for Multimodal Emotion
by: Zhu, Jichao, et al.
Published: (2026)
by: Zhu, Jichao, et al.
Published: (2026)
Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
by: Luo, Jie, et al.
Published: (2025)
by: Luo, Jie, et al.
Published: (2025)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation
by: Ma, Juncheng, et al.
Published: (2024)
by: Ma, Juncheng, et al.
Published: (2024)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
Unsupervised Instance Segmentation with Superpixels
by: Hoang, Cuong Manh
Published: (2025)
by: Hoang, Cuong Manh
Published: (2025)
Federated Unsupervised Semantic Segmentation
by: Charalampakis, Evangelos, et al.
Published: (2025)
by: Charalampakis, Evangelos, et al.
Published: (2025)
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
by: Zheng, Yuze, et al.
Published: (2024)
by: Zheng, Yuze, et al.
Published: (2024)
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
Single Image, Any Face: Generalisable 3D Face Generation
by: Wang, Wenqing, et al.
Published: (2024)
by: Wang, Wenqing, et al.
Published: (2024)
Dynamic Avatar-Scene Rendering from Human-centric Context
by: Wang, Wenqing, et al.
Published: (2025)
by: Wang, Wenqing, et al.
Published: (2025)
Unsupervised Synthetic Image Attribution: Alignment and Disentanglement
by: Liu, Zongfang, et al.
Published: (2026)
by: Liu, Zongfang, et al.
Published: (2026)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
by: Yu, Yinfeng, et al.
Published: (2025)
by: Yu, Yinfeng, et al.
Published: (2025)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
by: Lee, Yujian, et al.
Published: (2026)
by: Lee, Yujian, et al.
Published: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Dynamic Modality-Camera Invariant Clustering for Unsupervised Visible-Infrared Person Re-identification
by: Yang, Yiming, et al.
Published: (2024)
by: Yang, Yiming, et al.
Published: (2024)
Adversarially Domain-adaptive Latent Diffusion for Unsupervised Semantic Segmentation
by: Yu, Jongmin, et al.
Published: (2024)
by: Yu, Jongmin, et al.
Published: (2024)
SeeingSounds: Learning Audio-to-Visual Alignment via Text
by: Carnemolla, Simone, et al.
Published: (2025)
by: Carnemolla, Simone, et al.
Published: (2025)
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts
by: Zhu, Chenyang, et al.
Published: (2024)
by: Zhu, Chenyang, et al.
Published: (2024)
Context-Gated Cross-Modal Perception with Visual Mamba for PET-CT Lung Tumor Segmentation
by: Ayllón, Elena Mulero, et al.
Published: (2025)
by: Ayllón, Elena Mulero, et al.
Published: (2025)
Similar Items
-
AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
by: Bhosale, Swapnil, et al.
Published: (2024) -
3D Audio-Visual Segmentation
by: Sokolov, Artem, et al.
Published: (2024) -
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024) -
Recognize Any Regions
by: Yang, Haosen, et al.
Published: (2023) -
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)