Audio Visual Segmentation Through Text Embeddings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Kyungbok, Zhang, You, Duan, Zhiyao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
SeeingSounds: Learning Audio-to-Visual Alignment via Text
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2025)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
Audio-Visual Segmentation via Unlabeled Frame Exploitation
von: Liu, Jinxiang, et al.
Veröffentlicht: (2024)
von: Liu, Jinxiang, et al.
Veröffentlicht: (2024)
PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation
von: Ling, Jun, et al.
Veröffentlicht: (2024)
von: Ling, Jun, et al.
Veröffentlicht: (2024)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
SAFIRE: Segment Any Forged Image Region
von: Kwon, Myung-Joon, et al.
Veröffentlicht: (2024)
von: Kwon, Myung-Joon, et al.
Veröffentlicht: (2024)
Tracking and Segmenting Anything in Any Modality
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
von: Zhang, Tianlu, et al.
Veröffentlicht: (2025)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
von: Tian, Zeyue, et al.
Veröffentlicht: (2026)
Audio-visual Event Localization on Portrait Mode Short Videos
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
von: Liu, Wuyang, et al.
Veröffentlicht: (2025)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
von: Shin, Yosub, et al.
Veröffentlicht: (2025)
Style-Preserving Lip Sync via Audio-Aware Style Reference
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
Text-Only Data Synthesis for Vision Language Model Training
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2025)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
von: Huang, Feizhen, et al.
Veröffentlicht: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
von: Qi, Xingqun, et al.
Veröffentlicht: (2023)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
von: Liu, Che, et al.
Veröffentlicht: (2026)
von: Liu, Che, et al.
Veröffentlicht: (2026)
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
von: Hei, Nailei, et al.
Veröffentlicht: (2024)
von: Hei, Nailei, et al.
Veröffentlicht: (2024)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
von: S, Sridhar, et al.
Veröffentlicht: (2025)
von: S, Sridhar, et al.
Veröffentlicht: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework
von: Arkhipkin, Vladimir, et al.
Veröffentlicht: (2024)
von: Arkhipkin, Vladimir, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024) -
Taming Modality Entanglement in Continual Audio-Visual Segmentation
von: Hong, Yuyang, et al.
Veröffentlicht: (2025) -
SeeingSounds: Learning Audio-to-Visual Alignment via Text
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025) -
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024) -
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)