Unleash the Potential of CLIP for Video Highlight Detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Han, Donghoon, Seo, Seunghyeon, Park, Eunhwan, Nam, Seong-Uk, Kwak, Nojun |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
par: Han, Donghoon, et autres
Publié: (2026)
par: Han, Donghoon, et autres
Publié: (2026)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
par: Han, Donghoon, et autres
Publié: (2024)
par: Han, Donghoon, et autres
Publié: (2024)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
par: Han, Donghoon, et autres
Publié: (2023)
par: Han, Donghoon, et autres
Publié: (2023)
Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain Generalization
par: Lee, Dongkwan, et autres
Publié: (2025)
par: Lee, Dongkwan, et autres
Publié: (2025)
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
par: Han, Sangyu, et autres
Publié: (2024)
par: Han, Sangyu, et autres
Publié: (2024)
Causal Interpretation of Sparse Autoencoder Features in Vision
par: Han, Sangyu, et autres
Publié: (2025)
par: Han, Sangyu, et autres
Publié: (2025)
Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)
par: Kim, Yearim, et autres
Publié: (2024)
par: Kim, Yearim, et autres
Publié: (2024)
PedaCo-Gen: Scaffolding Pedagogical Agency in Human-AI Collaborative Video Authoring
par: Baek, Injun, et autres
Publié: (2026)
par: Baek, Injun, et autres
Publié: (2026)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
par: Um, Sung Jin, et autres
Publié: (2025)
par: Um, Sung Jin, et autres
Publié: (2025)
DivCon-NeRF: Diverse and Consistent Ray Augmentation for Few-Shot NeRF
par: Lee, Ingyun, et autres
Publié: (2025)
par: Lee, Ingyun, et autres
Publié: (2025)
MSG Score: Automated Video Verification for Reliable Multi-Scene Generation
par: Yoon, Daewon, et autres
Publié: (2024)
par: Yoon, Daewon, et autres
Publié: (2024)
LoGoColor: Local-Global 3D Colorization for 360° Scenes
par: Chang, Yeonjin, et autres
Publié: (2025)
par: Chang, Yeonjin, et autres
Publié: (2025)
The Role of Teacher Calibration in Knowledge Distillation
par: Kim, Suyoung, et autres
Publié: (2025)
par: Kim, Suyoung, et autres
Publié: (2025)
Real-Time Intuitive AI Drawing System for Collaboration: Enhancing Human Creativity through Formal and Contextual Intent Integration
par: Song, Jookyung, et autres
Publié: (2025)
par: Song, Jookyung, et autres
Publié: (2025)
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
par: Lee, Junhoo, et autres
Publié: (2026)
par: Lee, Junhoo, et autres
Publié: (2026)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
par: Choi, Hahyeon, et autres
Publié: (2025)
par: Choi, Hahyeon, et autres
Publié: (2025)
ROODI: Reconstructing Occluded Objects with Denoising Inpainters
par: Chang, Yeonjin, et autres
Publié: (2025)
par: Chang, Yeonjin, et autres
Publié: (2025)
ARC-NeRF: Area Ray Casting for Broader Unseen View Coverage in Few-shot Object Rendering
par: Seo, Seunghyeon, et autres
Publié: (2024)
par: Seo, Seunghyeon, et autres
Publié: (2024)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
par: Zeng, Gangyan, et autres
Publié: (2024)
par: Zeng, Gangyan, et autres
Publié: (2024)
S3D: Sketch-Driven 3D Model Generation
par: Song, Hail, et autres
Publié: (2025)
par: Song, Hail, et autres
Publié: (2025)
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
par: Kim, Suyoung, et autres
Publié: (2026)
par: Kim, Suyoung, et autres
Publié: (2026)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
par: Park, Kyu Ri, et autres
Publié: (2025)
par: Park, Kyu Ri, et autres
Publié: (2025)
Edge-case Synthesis for Fisheye Object Detection: A Data-centric Perspective
par: Kim, Seunghyeon, et autres
Publié: (2025)
par: Kim, Seunghyeon, et autres
Publié: (2025)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
par: Barbakos, Spyros, et autres
Publié: (2025)
par: Barbakos, Spyros, et autres
Publié: (2025)
The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP
par: Nam, Kahyeon, et autres
Publié: (2026)
par: Nam, Kahyeon, et autres
Publié: (2026)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
par: Paul, Dhiman, et autres
Publié: (2024)
par: Paul, Dhiman, et autres
Publié: (2024)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
par: Kim, Nayeong, et autres
Publié: (2025)
par: Kim, Nayeong, et autres
Publié: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
par: Gao, Bin-Bin, et autres
Publié: (2025)
par: Gao, Bin-Bin, et autres
Publié: (2025)
Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation
par: Lee, ByeongCheol, et autres
Publié: (2026)
par: Lee, ByeongCheol, et autres
Publié: (2026)
CLIP Unreasonable Potential in Single-Shot Face Recognition
par: Luu, Nhan T.
Publié: (2024)
par: Luu, Nhan T.
Publié: (2024)
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
par: Yu, Jiaao, et autres
Publié: (2025)
par: Yu, Jiaao, et autres
Publié: (2025)
Automated Detection of Sport Highlights from Audio and Video Sources
par: Della Santa, Francesco, et autres
Publié: (2025)
par: Della Santa, Francesco, et autres
Publié: (2025)
Generating Narrated Lecture Videos from Slides with Synchronized Highlights
par: Holmberg, Alexander
Publié: (2025)
par: Holmberg, Alexander
Publié: (2025)
HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting
par: Lee, Jeongeun, et autres
Publié: (2025)
par: Lee, Jeongeun, et autres
Publié: (2025)
Unleashing the Power of CNN and Transformer for Balanced RGB-Event Video Recognition
par: Wang, Xiao, et autres
Publié: (2023)
par: Wang, Xiao, et autres
Publié: (2023)
Illuminating Salient Contributions in Neuron Activation with Attribution Equilibrium
par: Nam, Woo-Jeoung, et autres
Publié: (2022)
par: Nam, Woo-Jeoung, et autres
Publié: (2022)
Video CLIP Model for Multi-View Echocardiography Interpretation
par: Takizawa, Ryo, et autres
Publié: (2025)
par: Takizawa, Ryo, et autres
Publié: (2025)
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
par: Ma, Wenxin, et autres
Publié: (2025)
par: Ma, Wenxin, et autres
Publié: (2025)
FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection
par: Chen, Yulin, et autres
Publié: (2025)
par: Chen, Yulin, et autres
Publié: (2025)
Source-Free Cross-Modal Knowledge Transfer by Unleashing the Potential of Task-Irrelevant Data
par: Zhu, Jinjing, et autres
Publié: (2024)
par: Zhu, Jinjing, et autres
Publié: (2024)
Documents similaires
-
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
par: Han, Donghoon, et autres
Publié: (2026) -
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
par: Han, Donghoon, et autres
Publié: (2024) -
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
par: Han, Donghoon, et autres
Publié: (2023) -
Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain Generalization
par: Lee, Dongkwan, et autres
Publié: (2025) -
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
par: Han, Sangyu, et autres
Publié: (2024)