ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Ji-Hyeon, Kim, Ho-Joong, Lee, Seong-Whan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
by: Ju, Yeong-Joon, et al.
Published: (2024)
by: Ju, Yeong-Joon, et al.
Published: (2024)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation
by: Kim, Nayeon, et al.
Published: (2024)
by: Kim, Nayeon, et al.
Published: (2024)
CW-BASS: Confidence-Weighted Boundary-Aware Learning for Semi-Supervised Semantic Segmentation
by: Tarubinga, Ebenezer, et al.
Published: (2025)
by: Tarubinga, Ebenezer, et al.
Published: (2025)
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
See, Rank, and Filter: Important Word-Aware Clip Filtering via Scene Understanding for Moment Retrieval and Highlight Detection
by: Lee, YuEun, et al.
Published: (2025)
by: Lee, YuEun, et al.
Published: (2025)
DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer
by: Kim, Ho-Joong, et al.
Published: (2025)
by: Kim, Ho-Joong, et al.
Published: (2025)
Video-GPT via Next Clip Diffusion
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression
by: Kim, Ho-Joong, et al.
Published: (2024)
by: Kim, Ho-Joong, et al.
Published: (2024)
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
by: Kang, Choongwon, et al.
Published: (2026)
by: Kang, Choongwon, et al.
Published: (2026)
Local Representative Token Guided Merging for Text-to-Image Generation
by: Lee, Min-Jeong, et al.
Published: (2025)
by: Lee, Min-Jeong, et al.
Published: (2025)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025)
by: Park, Geon, et al.
Published: (2025)
DQE-CIR: Distinctive Query Embeddings through Learnable Attribute Weights and Target Relative Negative Sampling in Composed Image Retrieval
by: Park, Geon, et al.
Published: (2026)
by: Park, Geon, et al.
Published: (2026)
MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
by: Park, Seojeong, et al.
Published: (2024)
by: Park, Seojeong, et al.
Published: (2024)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
by: Woo, Young Beom, et al.
Published: (2025)
by: Woo, Young Beom, et al.
Published: (2025)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
by: Truong, Quang-Trung, et al.
Published: (2025)
by: Truong, Quang-Trung, et al.
Published: (2025)
ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation
by: Li, Shengze, et al.
Published: (2024)
by: Li, Shengze, et al.
Published: (2024)
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
by: Tan, Wenhui, et al.
Published: (2026)
by: Tan, Wenhui, et al.
Published: (2026)
LUMINA-Net: Low-light Upgrade through Multi-stage Illumination and Noise Adaptation Network for Image Enhancement
by: Siddiqua, Namrah, et al.
Published: (2025)
by: Siddiqua, Namrah, et al.
Published: (2025)
Illuminating Salient Contributions in Neuron Activation with Attribution Equilibrium
by: Nam, Woo-Jeoung, et al.
Published: (2022)
by: Nam, Woo-Jeoung, et al.
Published: (2022)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
Objectomaly: Objectness-Aware Refinement for OoD Segmentation with Structural Consistency and Boundary Precision
by: Song, Jeonghoon, et al.
Published: (2025)
by: Song, Jeonghoon, et al.
Published: (2025)
GradMix: Gradient-based Selective Mixup for Robust Data Augmentation in Class-Incremental Learning
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
by: Jung, Gunho, et al.
Published: (2025)
by: Jung, Gunho, et al.
Published: (2025)
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
by: Jing, Xiaolun, et al.
Published: (2024)
by: Jing, Xiaolun, et al.
Published: (2024)
Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution Guidance
by: Hong, Jung-Ho, et al.
Published: (2023)
by: Hong, Jung-Ho, et al.
Published: (2023)
FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization
by: Kim, Seung-Wook, et al.
Published: (2025)
by: Kim, Seung-Wook, et al.
Published: (2025)
Reasoning over the Behaviour of Objects in Video-Clips for Adverb-Type Recognition
by: Seshadri, Amrit Diggavi, et al.
Published: (2023)
by: Seshadri, Amrit Diggavi, et al.
Published: (2023)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
by: Hong, Jung-Ho, et al.
Published: (2025)
by: Hong, Jung-Ho, et al.
Published: (2025)
Explaining generative diffusion models via visual analysis for interpretable decision-making process
by: Park, Ji-Hoon, et al.
Published: (2024)
by: Park, Ji-Hoon, et al.
Published: (2024)
Generic Event Boundary Detection via Denoising Diffusion
by: Hwang, Jaejun, et al.
Published: (2025)
by: Hwang, Jaejun, et al.
Published: (2025)
Enhancing Contrastive Learning with Efficient Combinatorial Positive Pairing
by: Kim, Jaeill, et al.
Published: (2024)
by: Kim, Jaeill, et al.
Published: (2024)
EEG-based Multimodal Representation Learning for Emotion Recognition
by: Yin, Kang, et al.
Published: (2024)
by: Yin, Kang, et al.
Published: (2024)
GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval
by: Zhang, Shihang, et al.
Published: (2026)
by: Zhang, Shihang, et al.
Published: (2026)
Transformer-based Clipped Contrastive Quantization Learning for Unsupervised Image Retrieval
by: Dubey, Ayush, et al.
Published: (2024)
by: Dubey, Ayush, et al.
Published: (2024)
NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
by: Gong, Zixuan, et al.
Published: (2024)
by: Gong, Zixuan, et al.
Published: (2024)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
ClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection
by: Lu, Hong, et al.
Published: (2025)
by: Lu, Hong, et al.
Published: (2025)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
Language-Guided Invariance Probing of Vision-Language Models
by: Lee, Jae Joong
Published: (2025)
by: Lee, Jae Joong
Published: (2025)
Similar Items
-
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
by: Ju, Yeong-Joon, et al.
Published: (2024) -
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025) -
Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation
by: Kim, Nayeon, et al.
Published: (2024) -
CW-BASS: Confidence-Weighted Boundary-Aware Learning for Semi-Supervised Semantic Segmentation
by: Tarubinga, Ebenezer, et al.
Published: (2025) -
In-Context Learning with Unpaired Clips for Instruction-based Video Editing
by: Liao, Xinyao, et al.
Published: (2025)