Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Jiaming, Li, Hanjun, Lin, Kun-Yu, Liang, Junwei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Human-Centric Transformer for Domain Adaptive Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning
by: Du, Jia-Run, et al.
Published: (2022)
by: Du, Jia-Run, et al.
Published: (2022)
Hierarchical Action Learning for Weakly-Supervised Action Segmentation
by: Huang, Junxian, et al.
Published: (2026)
by: Huang, Junxian, et al.
Published: (2026)
VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
by: Wang, Zhuming, et al.
Published: (2025)
by: Wang, Zhuming, et al.
Published: (2025)
EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations
by: Liu, Jiayi, et al.
Published: (2025)
by: Liu, Jiayi, et al.
Published: (2025)
Enhancing Weakly Supervised Semantic Segmentation with Multi-modal Foundation Models: An End-to-End Approach
by: Ravanbakhsh, Elham, et al.
Published: (2024)
by: Ravanbakhsh, Elham, et al.
Published: (2024)
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
by: Zhou, Jiaming, et al.
Published: (2025)
by: Zhou, Jiaming, et al.
Published: (2025)
End2end-ALARA: Approaching the ALARA Law in CT Imaging with End-to-end Learning
by: Tao, Xi, et al.
Published: (2025)
by: Tao, Xi, et al.
Published: (2025)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
by: Zhou, Milton, et al.
Published: (2026)
by: Zhou, Milton, et al.
Published: (2026)
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
by: Zhou, Xingcheng, et al.
Published: (2025)
by: Zhou, Xingcheng, et al.
Published: (2025)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Uncovering the Handwritten Text in the Margins: End-to-end Handwritten Text Detection and Recognition
by: Cheng, Liang, et al.
Published: (2023)
by: Cheng, Liang, et al.
Published: (2023)
An Effective End-to-End Solution for Multimodal Action Recognition
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
2by2: Weakly-Supervised Learning for Global Action Segmentation
by: Bueno-Benito, Elena, et al.
Published: (2024)
by: Bueno-Benito, Elena, et al.
Published: (2024)
MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild
by: Fang, Xi, et al.
Published: (2024)
by: Fang, Xi, et al.
Published: (2024)
Partial Weakly-Supervised Oriented Object Detection
by: Liu, Mingxin, et al.
Published: (2025)
by: Liu, Mingxin, et al.
Published: (2025)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
Pose-Aware Weakly-Supervised Action Segmentation
by: Zhao, Seth Z., et al.
Published: (2025)
by: Zhao, Seth Z., et al.
Published: (2025)
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
by: Luo, Dongliang, et al.
Published: (2025)
by: Luo, Dongliang, et al.
Published: (2025)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
Towards Adaptive Pseudo-label Learning for Semi-Supervised Temporal Action Localization
by: Zhou, Feixiang, et al.
Published: (2024)
by: Zhou, Feixiang, et al.
Published: (2024)
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
by: Sun, Baoli, et al.
Published: (2025)
by: Sun, Baoli, et al.
Published: (2025)
CloudMatch: Weak-to-Strong Consistency Learning for Semi-Supervised Cloud Detection
by: Zhao, Jiayi, et al.
Published: (2026)
by: Zhao, Jiayi, et al.
Published: (2026)
Complete Instances Mining for Weakly Supervised Instance Segmentation
by: Li, Zecheng, et al.
Published: (2024)
by: Li, Zecheng, et al.
Published: (2024)
Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
by: Tang, Wenhao, et al.
Published: (2025)
by: Tang, Wenhao, et al.
Published: (2025)
MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
by: Strohmeyer, Tim, et al.
Published: (2026)
by: Strohmeyer, Tim, et al.
Published: (2026)
WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
WP-CrackNet: A Collaborative Adversarial Learning Framework for End-to-End Weakly-Supervised Road Crack Detection
by: Ma, Nachuan, et al.
Published: (2025)
by: Ma, Nachuan, et al.
Published: (2025)
Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation
by: Ma, Enhui, et al.
Published: (2024)
by: Ma, Enhui, et al.
Published: (2024)
Boosting Semi-Supervised Temporal Action Localization by Learning from Non-Target Classes
by: Xia, Kun, et al.
Published: (2024)
by: Xia, Kun, et al.
Published: (2024)
Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise Correction
by: Zhang, Quan, et al.
Published: (2025)
by: Zhang, Quan, et al.
Published: (2025)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
Weakly Supervised Object Detection for Automatic Tooth-marked Tongue Recognition
by: Zhang, Yongcun, et al.
Published: (2024)
by: Zhang, Yongcun, et al.
Published: (2024)
Task-Specific Distance Correlation Matching for Few-Shot Action Recognition
by: Long, Fei, et al.
Published: (2025)
by: Long, Fei, et al.
Published: (2025)
Similar Items
-
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024) -
Human-Centric Transformer for Domain Adaptive Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024) -
Weakly-Supervised Temporal Action Localization by Progressive Complementary Learning
by: Du, Jia-Run, et al.
Published: (2022) -
Hierarchical Action Learning for Weakly-Supervised Action Segmentation
by: Huang, Junxian, et al.
Published: (2026) -
VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
by: Wang, Zhuming, et al.
Published: (2025)