Audio-visual Event Localization on Portrait Mode Short Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Wuyang, Chai, Yi, Yan, Yongpeng, Ren, Yanzhen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
by: Tan, Weiting, et al.
Published: (2026)
by: Tan, Weiting, et al.
Published: (2026)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
Question-Answering Dense Video Events
by: Qin, Hangyu, et al.
Published: (2024)
by: Qin, Hangyu, et al.
Published: (2024)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
by: Qu, Qiang, et al.
Published: (2025)
by: Qu, Qiang, et al.
Published: (2025)
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
by: Shin, Yosub, et al.
Published: (2025)
by: Shin, Yosub, et al.
Published: (2025)
Do Joint Audio-Video Generation Models Understand Physics?
by: Cui, Zijun, et al.
Published: (2026)
by: Cui, Zijun, et al.
Published: (2026)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
by: Huang, Feizhen, et al.
Published: (2025)
by: Huang, Feizhen, et al.
Published: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
Diffusion Models for Joint Audio-Video Generation
by: La Torre, Alejandro Paredes
Published: (2026)
by: La Torre, Alejandro Paredes
Published: (2026)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
by: Gu, Jing, et al.
Published: (2024)
by: Gu, Jing, et al.
Published: (2024)
Learning to Generate Conditional Tri-plane for 3D-aware Expression Controllable Portrait Animation
by: Ki, Taekyung, et al.
Published: (2024)
by: Ki, Taekyung, et al.
Published: (2024)
Audio-Guided Visual Perception for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Apollo: Unified Multi-Task Audio-Video Joint Generation
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
by: Ku, Max, et al.
Published: (2024)
by: Ku, Max, et al.
Published: (2024)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
by: Yang, Jialiang, et al.
Published: (2026)
by: Yang, Jialiang, et al.
Published: (2026)
Decoupled Audio-Visual Dataset Distillation
by: Li, Wenyuan, et al.
Published: (2025)
by: Li, Wenyuan, et al.
Published: (2025)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
Audio Visual Segmentation Through Text Embeddings
by: Lee, Kyungbok, et al.
Published: (2025)
by: Lee, Kyungbok, et al.
Published: (2025)
Bernini: Latent Semantic Planning for Video Diffusion
by: Bernini Team, et al.
Published: (2026)
by: Bernini Team, et al.
Published: (2026)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
by: Yu, Jiashuo, et al.
Published: (2025)
by: Yu, Jiashuo, et al.
Published: (2025)
EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation
by: Qi, Xingqun, et al.
Published: (2023)
by: Qi, Xingqun, et al.
Published: (2023)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
by: Zeng, Xiangyu, et al.
Published: (2024)
by: Zeng, Xiangyu, et al.
Published: (2024)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
by: Tian, Zeyue, et al.
Published: (2026)
by: Tian, Zeyue, et al.
Published: (2026)
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
by: Qu, Qiang, et al.
Published: (2024)
by: Qu, Qiang, et al.
Published: (2024)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition
by: Yu, Zongyou, et al.
Published: (2024)
by: Yu, Zongyou, et al.
Published: (2024)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
by: Marinoni, Christian, et al.
Published: (2025)
by: Marinoni, Christian, et al.
Published: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
Style-Preserving Lip Sync via Audio-Aware Style Reference
by: Zhong, Weizhi, et al.
Published: (2024)
by: Zhong, Weizhi, et al.
Published: (2024)
Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
by: Liu, Jiajun, et al.
Published: (2024)
by: Liu, Jiajun, et al.
Published: (2024)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
by: Li, Guangyao, et al.
Published: (2024)
by: Li, Guangyao, et al.
Published: (2024)
Video Seal: Open and Efficient Video Watermarking
by: Fernandez, Pierre, et al.
Published: (2024)
by: Fernandez, Pierre, et al.
Published: (2024)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
by: Yuan, Hangjie, et al.
Published: (2025)
by: Yuan, Hangjie, et al.
Published: (2025)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
by: Cai, Dongnuan, et al.
Published: (2026)
by: Cai, Dongnuan, et al.
Published: (2026)
Similar Items
-
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
by: Tan, Weiting, et al.
Published: (2026) -
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
by: Zhou, Jinxing, et al.
Published: (2024) -
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025) -
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025) -
Question-Answering Dense Video Events
by: Qin, Hangyu, et al.
Published: (2024)