GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Bowen, Cao, Yun, He, Chen, Su, Xiaosu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
von: Deng, Yuchen, et al.
Veröffentlicht: (2025)
von: Deng, Yuchen, et al.
Veröffentlicht: (2025)
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
von: Hur, Chan, et al.
Veröffentlicht: (2025)
von: Hur, Chan, et al.
Veröffentlicht: (2025)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
von: Han, Gyuwon, et al.
Veröffentlicht: (2026)
von: Han, Gyuwon, et al.
Veröffentlicht: (2026)
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
von: Zhao, Zecheng, et al.
Veröffentlicht: (2025)
von: Zhao, Zecheng, et al.
Veröffentlicht: (2025)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
von: Yang, Shaoshu, et al.
Veröffentlicht: (2025)
von: Yang, Shaoshu, et al.
Veröffentlicht: (2025)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
Uncertainty-Gated Region-Level Retrieval for Robust Semantic Segmentation
von: Rajan, Shreshth, et al.
Veröffentlicht: (2025)
von: Rajan, Shreshth, et al.
Veröffentlicht: (2025)
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
von: Liu, Yunze, et al.
Veröffentlicht: (2026)
von: Liu, Yunze, et al.
Veröffentlicht: (2026)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
von: Yu, Li, et al.
Veröffentlicht: (2025)
von: Yu, Li, et al.
Veröffentlicht: (2025)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
von: Liu, Weijia, et al.
Veröffentlicht: (2025)
von: Liu, Weijia, et al.
Veröffentlicht: (2025)
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
von: Ding, Xin, et al.
Veröffentlicht: (2025)
von: Ding, Xin, et al.
Veröffentlicht: (2025)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
von: Jin, Xiaojie, et al.
Veröffentlicht: (2023)
von: Jin, Xiaojie, et al.
Veröffentlicht: (2023)
Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
Text-Video Multi-Grained Integration for Video Moment Montage
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
von: Yin, Zhihui, et al.
Veröffentlicht: (2024)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning
von: Kim, Seongah, et al.
Veröffentlicht: (2026)
von: Kim, Seongah, et al.
Veröffentlicht: (2026)
TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
von: Feng, Yue, et al.
Veröffentlicht: (2025)
von: Feng, Yue, et al.
Veröffentlicht: (2025)
Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
von: Zhang, Jiayu, et al.
Veröffentlicht: (2026)
von: Zhang, Jiayu, et al.
Veröffentlicht: (2026)
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
von: Zhao, Zecheng, et al.
Veröffentlicht: (2026)
von: Zhao, Zecheng, et al.
Veröffentlicht: (2026)
Semantic Audio-Visual Navigation in Continuous Environments
von: Zeng, Yichen, et al.
Veröffentlicht: (2026)
von: Zeng, Yichen, et al.
Veröffentlicht: (2026)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
Fine-Grained Zero-Shot Composed Image Retrieval with Complementary Visual-Semantic Integration
von: Ye, Yongcong, et al.
Veröffentlicht: (2026)
von: Ye, Yongcong, et al.
Veröffentlicht: (2026)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
von: Cai, Zhixi, et al.
Veröffentlicht: (2025)
von: Cai, Zhixi, et al.
Veröffentlicht: (2025)
Semantic Frame Interpolation
von: Hong, Yijia, et al.
Veröffentlicht: (2025)
von: Hong, Yijia, et al.
Veröffentlicht: (2025)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
von: Yun, Heeseung, et al.
Veröffentlicht: (2024)
von: Yun, Heeseung, et al.
Veröffentlicht: (2024)
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating
von: Cao, Xiangkui, et al.
Veröffentlicht: (2026)
von: Cao, Xiangkui, et al.
Veröffentlicht: (2026)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
von: Li, Jungang, et al.
Veröffentlicht: (2024)
von: Li, Jungang, et al.
Veröffentlicht: (2024)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
von: Luo, Dezhao, et al.
Veröffentlicht: (2024)
von: Luo, Dezhao, et al.
Veröffentlicht: (2024)
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
von: Cao, Zhe, et al.
Veröffentlicht: (2025)
von: Cao, Zhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
von: Jeong, Boseung, et al.
Veröffentlicht: (2025) -
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
von: Deng, Yuchen, et al.
Veröffentlicht: (2025) -
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
von: Hur, Chan, et al.
Veröffentlicht: (2025) -
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
von: Han, Gyuwon, et al.
Veröffentlicht: (2026) -
Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
von: Zhao, Zecheng, et al.
Veröffentlicht: (2025)