Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Lim, Geuntaek, Kim, Hyunwoo, Kim, Joonsoo, Choi, Yukyung |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video-Oasis: Rethinking Evaluation of Video Understanding
by: Lim, Geuntaek, et al.
Published: (2026)
by: Lim, Geuntaek, et al.
Published: (2026)
Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning
by: Du, Jia-Run, et al.
Published: (2022)
by: Du, Jia-Run, et al.
Published: (2022)
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
by: Ma, Yunchuan, et al.
Published: (2026)
by: Ma, Yunchuan, et al.
Published: (2026)
MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splatting
by: Kwak, Sangwoon, et al.
Published: (2025)
by: Kwak, Sangwoon, et al.
Published: (2025)
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
by: Lee, Sangin, et al.
Published: (2026)
by: Lee, Sangin, et al.
Published: (2026)
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
by: Su, Rui, et al.
Published: (2025)
by: Su, Rui, et al.
Published: (2025)
ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
by: Park, Jihwan, et al.
Published: (2025)
by: Park, Jihwan, et al.
Published: (2025)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models
by: Zhang, Quan, et al.
Published: (2024)
by: Zhang, Quan, et al.
Published: (2024)
RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
by: Park, Jihwan, et al.
Published: (2026)
by: Park, Jihwan, et al.
Published: (2026)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
Multi-Modal Guided Multi-Source Domain Adaptation for Object Detection
by: Lee, Sangin, et al.
Published: (2026)
by: Lee, Sangin, et al.
Published: (2026)
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026)
by: Wang, Fengshun, et al.
Published: (2026)
Weakly Supervised Video Scene Graph Generation via Natural Language Supervision
by: Kim, Kibum, et al.
Published: (2025)
by: Kim, Kibum, et al.
Published: (2025)
Efficient multi-view training for 3D Gaussian Splatting
by: Choi, Minhyuk, et al.
Published: (2025)
by: Choi, Minhyuk, et al.
Published: (2025)
Representation Shift: Unifying Token Compression with FlashAttention
by: Choi, Joonmyung, et al.
Published: (2025)
by: Choi, Joonmyung, et al.
Published: (2025)
Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise Correction
by: Zhang, Quan, et al.
Published: (2025)
by: Zhang, Quan, et al.
Published: (2025)
Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormer
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
TRiGS: Temporal Rigid-Body Motion for Scalable 4D Gaussian Splatting
by: Yeom, Suwoong, et al.
Published: (2026)
by: Yeom, Suwoong, et al.
Published: (2026)
Boosting Cross-spectral Unsupervised Domain Adaptation for Thermal Semantic Segmentation
by: Kwon, Seokjun, et al.
Published: (2025)
by: Kwon, Seokjun, et al.
Published: (2025)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2026)
by: Ko, Dohwan, et al.
Published: (2026)
Online Temporal Action Localization with Memory-Augmented Transformer
by: Song, Youngkil, et al.
Published: (2024)
by: Song, Youngkil, et al.
Published: (2024)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
by: Bae, Kyungho, et al.
Published: (2025)
by: Bae, Kyungho, et al.
Published: (2025)
Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning
by: Xu, Wenbo, et al.
Published: (2025)
by: Xu, Wenbo, et al.
Published: (2025)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
by: Choi, Dasol, et al.
Published: (2026)
by: Choi, Dasol, et al.
Published: (2026)
Rethinking Saliency-Guided Weakly-Supervised Semantic Segmentation
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
WS-IMUBench: Can Weakly Supervised Methods from Audio, Image, and Video Be Adapted for IMU-based Temporal Action Localization?
by: Li, Pei, et al.
Published: (2026)
by: Li, Pei, et al.
Published: (2026)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
by: Bae, Kyungho, et al.
Published: (2023)
by: Bae, Kyungho, et al.
Published: (2023)
Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
by: Choi, Seung hee, et al.
Published: (2026)
by: Choi, Seung hee, et al.
Published: (2026)
Hierarchical Action Learning for Weakly-Supervised Action Segmentation
by: Huang, Junxian, et al.
Published: (2026)
by: Huang, Junxian, et al.
Published: (2026)
Retinal Layer Segmentation in OCT Images With 2.5D Cross-slice Feature Fusion Module for Glaucoma Assessment
by: Kim, Hyunwoo, et al.
Published: (2026)
by: Kim, Hyunwoo, et al.
Published: (2026)
A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization
by: Xu, Wenbo, et al.
Published: (2025)
by: Xu, Wenbo, et al.
Published: (2025)
Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
by: Han, Cailing, et al.
Published: (2026)
by: Han, Cailing, et al.
Published: (2026)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
by: Jin, In-Hwan, et al.
Published: (2025)
by: Jin, In-Hwan, et al.
Published: (2025)
Similar Items
-
Video-Oasis: Rethinking Evaluation of Video Understanding
by: Lim, Geuntaek, et al.
Published: (2026) -
Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning
by: Du, Jia-Run, et al.
Published: (2022) -
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
by: Ma, Yunchuan, et al.
Published: (2026) -
MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splatting
by: Kwak, Sangwoon, et al.
Published: (2025) -
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
by: Lee, Sangin, et al.
Published: (2026)