STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jaewoo, Yoon, Jaehong, Kim, Wonjae, Kim, Yunji, Hwang, Sung Ju |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022)
by: Hwang, Sunil, et al.
Published: (2022)
Continual Learning: Forget-free Winning Subnetworks for Video Representations
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
by: Yeo, Woongyeong, et al.
Published: (2025)
by: Yeo, Woongyeong, et al.
Published: (2025)
Progressive Fourier Neural Representation for Sequential Video Compilation
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
Self-Refining Video Sampling
by: Jang, Sangwon, et al.
Published: (2026)
by: Jang, Sangwon, et al.
Published: (2026)
Concept-skill Transferability-based Data Selection for Large Vision-Language Models
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
by: Yoon, Junho, et al.
Published: (2025)
by: Yoon, Junho, et al.
Published: (2025)
Probabilistic Language-Image Pre-Training
by: Chun, Sanghyuk, et al.
Published: (2024)
by: Chun, Sanghyuk, et al.
Published: (2024)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
Progressive Local Alignment for Medical Multimodal Pre-training
by: Yan, Huimin, et al.
Published: (2025)
by: Yan, Huimin, et al.
Published: (2025)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)
by: Jun, Youngjun, et al.
Published: (2026)
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
by: Ki, Taekyung, et al.
Published: (2026)
by: Ki, Taekyung, et al.
Published: (2026)
Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
by: Kim, Jaihoon, et al.
Published: (2025)
by: Kim, Jaihoon, et al.
Published: (2025)
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)
by: Lee, Youngwan, et al.
Published: (2024)
ECLIPSE: Efficient Continual Learning in Panoptic Segmentation with Visual Prompt Tuning
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
CMTA: Cross-Modal Temporal Alignment for Event-guided Video Deblurring
by: Kim, Taewoo, et al.
Published: (2024)
by: Kim, Taewoo, et al.
Published: (2024)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
VideoMamba: Spatio-Temporal Selective State Space Model
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
by: Lee, Yeonkyung, et al.
Published: (2026)
by: Lee, Yeonkyung, et al.
Published: (2026)
FedPOD: the deployable units of training for federated learning
by: Kim, Daewoon, et al.
Published: (2025)
by: Kim, Daewoon, et al.
Published: (2025)
ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models
by: Sung, Yi-Lin, et al.
Published: (2023)
by: Sung, Yi-Lin, et al.
Published: (2023)
Rethinking Saliency-Guided Weakly-Supervised Semantic Segmentation
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
Feature Unlearning for Pre-trained GANs and VAEs
by: Moon, Saemi, et al.
Published: (2023)
by: Moon, Saemi, et al.
Published: (2023)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
by: Wang, Hengkang, et al.
Published: (2025)
by: Wang, Hengkang, et al.
Published: (2025)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Do Pre-trained Models Benefit Equally in Continual Learning?
by: Lee, Kuan-Ying, et al.
Published: (2022)
by: Lee, Kuan-Ying, et al.
Published: (2022)
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation
by: Min, Dongchan, et al.
Published: (2022)
by: Min, Dongchan, et al.
Published: (2022)
Object-aware Sound Source Localization via Audio-Visual Scene Understanding
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
by: Kim, Kangsan, et al.
Published: (2024)
by: Kim, Kangsan, et al.
Published: (2024)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
by: Kang, Ju Yeon, et al.
Published: (2025)
by: Kang, Ju Yeon, et al.
Published: (2025)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
by: Jang, Jinhyeok, et al.
Published: (2025)
by: Jang, Jinhyeok, et al.
Published: (2025)
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
by: So, Yerim, et al.
Published: (2026)
by: So, Yerim, et al.
Published: (2026)
Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
by: Seon, Juhyeong, et al.
Published: (2024)
by: Seon, Juhyeong, et al.
Published: (2024)
Video-Language Alignment via Spatio-Temporal Graph Transformer
by: Zhang, Shi-Xue, et al.
Published: (2024)
by: Zhang, Shi-Xue, et al.
Published: (2024)
Extract Free Dense Misalignment from CLIP
by: Nam, JeongYeon, et al.
Published: (2024)
by: Nam, JeongYeon, et al.
Published: (2024)
Similar Items
-
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
by: Lee, Daeun, et al.
Published: (2024) -
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022) -
Continual Learning: Forget-free Winning Subnetworks for Video Representations
by: Kang, Haeyong, et al.
Published: (2023) -
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
by: Yeo, Woongyeong, et al.
Published: (2025) -
Progressive Fourier Neural Representation for Sequential Video Compilation
by: Kang, Haeyong, et al.
Published: (2023)