Audio-Infused Automatic Image Colorization by Exploiting Audio Scene Semantics
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Pengcheng, Chen, Yanxiang, Zhao, Yang, Zhang, Zhao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
by: Zhao, Pengcheng, et al.
Published: (2024)
by: Zhao, Pengcheng, et al.
Published: (2024)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
Decoupling Semantics and Fingerprints: A Universal Representation for AI-Generated Image Detection
by: Wang, Zhiyuan, et al.
Published: (2026)
by: Wang, Zhiyuan, et al.
Published: (2026)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Cai, Dongnuan, et al.
Published: (2026)
by: Cai, Dongnuan, et al.
Published: (2026)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025)
by: Xu, Sicheng, et al.
Published: (2025)
Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation
by: Ma, Juncheng, et al.
Published: (2024)
by: Ma, Juncheng, et al.
Published: (2024)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)
by: Jiang, Yuanyuan, et al.
Published: (2022)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
by: Tian, Zeyue, et al.
Published: (2026)
by: Tian, Zeyue, et al.
Published: (2026)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
by: Zhang, Guohui, et al.
Published: (2026)
by: Zhang, Guohui, et al.
Published: (2026)
Unsupervised Audio-Visual Segmentation with Modality Alignment
by: Bhosale, Swapnil, et al.
Published: (2024)
by: Bhosale, Swapnil, et al.
Published: (2024)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
by: Lee, Yujian, et al.
Published: (2026)
by: Lee, Yujian, et al.
Published: (2026)
video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM
by: Sun, Guangzhi, et al.
Published: (2025)
by: Sun, Guangzhi, et al.
Published: (2025)
Quaternion Generative Adversarial Neural Networks and Applications to Color Image Inpainting
by: Wang, Duan, et al.
Published: (2024)
by: Wang, Duan, et al.
Published: (2024)
Comparative Analysis of Image, Video, and Audio Classifiers for Automated News Video Segmentation
by: Attard, Jonathan, et al.
Published: (2025)
by: Attard, Jonathan, et al.
Published: (2025)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
by: Li, Guangyao, et al.
Published: (2024)
by: Li, Guangyao, et al.
Published: (2024)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
by: Yang, Jialiang, et al.
Published: (2026)
by: Yang, Jialiang, et al.
Published: (2026)
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
LACOSTE: Exploiting stereo and temporal contexts for surgical instrument segmentation
by: Wang, Qiyuan, et al.
Published: (2024)
by: Wang, Qiyuan, et al.
Published: (2024)
Lightweight Remote Sensing Scene Classification on Edge Devices via Knowledge Distillation and Early-exit
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
Exploiting Motion Prior for Accurate Pose Estimation of Dashboard Cameras
by: Lu, Yipeng, et al.
Published: (2024)
by: Lu, Yipeng, et al.
Published: (2024)
Dynamic Multi-Target Fusion for Efficient Audio-Visual Navigation
by: Yu, Yinfeng, et al.
Published: (2025)
by: Yu, Yinfeng, et al.
Published: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
STNet: Deep Audio-Visual Fusion Network for Robust Speaker Tracking
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
MINT: Memory-Infused Prompt Tuning at Test-time for CLIP
by: Yi, Jiaming, et al.
Published: (2025)
by: Yi, Jiaming, et al.
Published: (2025)
IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer
by: Chen, Bo, et al.
Published: (2025)
by: Chen, Bo, et al.
Published: (2025)
Feature-Space Semantic Invariance: Enhanced OOD Detection for Open-Set Domain Generalization
by: Wang, Haoliang, et al.
Published: (2024)
by: Wang, Haoliang, et al.
Published: (2024)
INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
by: Zhu, Yongming, et al.
Published: (2024)
by: Zhu, Yongming, et al.
Published: (2024)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
Audio Visual Segmentation Through Text Embeddings
by: Lee, Kyungbok, et al.
Published: (2025)
by: Lee, Kyungbok, et al.
Published: (2025)
Exploiting Diffusion Prior for Out-of-Distribution Detection
by: Zhu, Armando, et al.
Published: (2024)
by: Zhu, Armando, et al.
Published: (2024)
TextMamba: Scene Text Detector with Mamba
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
Can Shape-Infused Joint Embeddings Improve Image-Conditioned 3D Diffusion?
by: Sbrolli, Cristian, et al.
Published: (2024)
by: Sbrolli, Cristian, et al.
Published: (2024)
Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
by: Liu, Delong, et al.
Published: (2023)
by: Liu, Delong, et al.
Published: (2023)
Similar Items
-
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
by: Zhao, Pengcheng, et al.
Published: (2024) -
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024) -
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
by: Yang, Bowen, et al.
Published: (2025) -
Decoupling Semantics and Fingerprints: A Universal Representation for AI-Generated Image Detection
by: Wang, Zhiyuan, et al.
Published: (2026) -
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Cai, Dongnuan, et al.
Published: (2026)