Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Yunchuan, Qing, Laiyun, Li, Guorong, Liu, Yuqing, Qi, Yuankai, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
SOVC: Subject-Oriented Video Captioning
von: Teng, Chang, et al.
Veröffentlicht: (2023)
von: Teng, Chang, et al.
Veröffentlicht: (2023)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
von: Tao, Zhuo, et al.
Veröffentlicht: (2025)
von: Tao, Zhuo, et al.
Veröffentlicht: (2025)
HR-Pro: Point-supervised Temporal Action Localization via Hierarchical Reliability Propagation
von: Zhang, Huaxin, et al.
Veröffentlicht: (2023)
von: Zhang, Huaxin, et al.
Veröffentlicht: (2023)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
Temporal Action Localization with Cross Layer Task Decoupling and Refinement
von: Li, Qiang, et al.
Veröffentlicht: (2024)
von: Li, Qiang, et al.
Veröffentlicht: (2024)
Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
von: Dai, Siran, et al.
Veröffentlicht: (2025)
von: Dai, Siran, et al.
Veröffentlicht: (2025)
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Decorrelating Structure via Adapters Makes Ensemble Learning Practical for Semi-supervised Learning
von: Wu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Wu, Jiaqi, et al.
Veröffentlicht: (2024)
Full-Stage Pseudo Label Quality Enhancement for Weakly-supervised Temporal Action Localization
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
CPR++: Object Localization via Single Coarse Point Supervision
von: Yu, Xuehui, et al.
Veröffentlicht: (2024)
von: Yu, Xuehui, et al.
Veröffentlicht: (2024)
Exploring Structural Degradation in Dense Representations for Self-supervised Learning
von: Dai, Siran, et al.
Veröffentlicht: (2025)
von: Dai, Siran, et al.
Veröffentlicht: (2025)
Convex Combination Consistency between Neighbors for Weakly-supervised Action Localization
von: Liu, Qinying, et al.
Veröffentlicht: (2022)
von: Liu, Qinying, et al.
Veröffentlicht: (2022)
Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
von: Yu, Xuan, et al.
Veröffentlicht: (2025)
von: Yu, Xuan, et al.
Veröffentlicht: (2025)
STAT: Towards Generalizable Temporal Action Localization
von: Liu, Yangcen, et al.
Veröffentlicht: (2024)
von: Liu, Yangcen, et al.
Veröffentlicht: (2024)
Text-Driven Diverse Facial Texture Generation via Progressive Latent-Space Refinement
von: Wang, Chi, et al.
Veröffentlicht: (2024)
von: Wang, Chi, et al.
Veröffentlicht: (2024)
Boosting Semi-Supervised Temporal Action Localization by Learning from Non-Target Classes
von: Xia, Kun, et al.
Veröffentlicht: (2024)
von: Xia, Kun, et al.
Veröffentlicht: (2024)
POTLoc: Pseudo-Label Oriented Transformer for Point-Supervised Temporal Action Localization
von: Vahdani, Elahe, et al.
Veröffentlicht: (2023)
von: Vahdani, Elahe, et al.
Veröffentlicht: (2023)
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
von: Chen, Yuxiao, et al.
Veröffentlicht: (2024)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2024)
Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance
von: Wang, Meng, et al.
Veröffentlicht: (2025)
von: Wang, Meng, et al.
Veröffentlicht: (2025)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
von: Cong, Gaoxiang, et al.
Veröffentlicht: (2026)
Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
von: Lv, Henglei, et al.
Veröffentlicht: (2024)
von: Lv, Henglei, et al.
Veröffentlicht: (2024)
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
von: Guo, Jiayi, et al.
Veröffentlicht: (2026)
von: Guo, Jiayi, et al.
Veröffentlicht: (2026)
Uncertainty-aware Long-tailed Weights Model the Utility of Pseudo-labels for Semi-supervised Learning
von: Wu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wu, Jiaqi, et al.
Veröffentlicht: (2025)
Enhancing Pre-trained Representation Classifiability can Boost its Interpretability
von: Shen, Shufan, et al.
Veröffentlicht: (2025)
von: Shen, Shufan, et al.
Veröffentlicht: (2025)
Coarse-to-Fine Monocular Re-Localization in OpenStreetMap via Semantic Alignment
von: Zou, Yuchen, et al.
Veröffentlicht: (2026)
von: Zou, Yuchen, et al.
Veröffentlicht: (2026)
From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point Cloud
von: Lu, Jingjing, et al.
Veröffentlicht: (2025)
von: Lu, Jingjing, et al.
Veröffentlicht: (2025)
MambaLCT: Boosting Tracking via Long-term Context State Space Model
von: Li, Xiaohai, et al.
Veröffentlicht: (2024)
von: Li, Xiaohai, et al.
Veröffentlicht: (2024)
Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
von: Zhu, Xinnan, et al.
Veröffentlicht: (2025)
von: Zhu, Xinnan, et al.
Veröffentlicht: (2025)
STAR-IOD: Scale-decoupled Topology Alignment with Pseudo-label Refinement for Remote Sensing Incremental Object Detection
von: Zhang, Yaoteng, et al.
Veröffentlicht: (2026)
von: Zhang, Yaoteng, et al.
Veröffentlicht: (2026)
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
von: Qi, Zhaobo, et al.
Veröffentlicht: (2024)
von: Qi, Zhaobo, et al.
Veröffentlicht: (2024)
PointSmile: Point Self-supervised Learning via Curriculum Mutual Information
von: Li, Xin, et al.
Veröffentlicht: (2023)
von: Li, Xin, et al.
Veröffentlicht: (2023)
FMI-TAL: Few-shot Multiple Instances Temporal Action Localization by Probability Distribution Learning and Interval Cluster Refinement
von: Wang, Fengshun, et al.
Veröffentlicht: (2024)
von: Wang, Fengshun, et al.
Veröffentlicht: (2024)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
von: Zhang, Mingfang, et al.
Veröffentlicht: (2025)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026) -
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024) -
SOVC: Subject-Oriented Video Captioning
von: Teng, Chang, et al.
Veröffentlicht: (2023) -
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
von: Zhao, Yiming, et al.
Veröffentlicht: (2025) -
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)