The Solution for Temporal Action Localisation Task of Perception Test Challenge 2024
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Yinan, Jiang, Qingyuan, Mei, Hongming, Yang, Yang, Tang, Jinhui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
by: Huang, Yurui, et al.
Published: (2024)
by: Huang, Yurui, et al.
Published: (2024)
The Solution for Single Object Tracking Task of Perception Test Challenge 2024
by: Zhong, Zhiqiang, et al.
Published: (2024)
by: Zhong, Zhiqiang, et al.
Published: (2024)
Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024
by: Ahn, Jinwoo, et al.
Published: (2024)
by: Ahn, Jinwoo, et al.
Published: (2024)
Cross-Task Attack: A Self-Supervision Generative Framework Based on Attention Shift
by: Zeng, Qingyuan, et al.
Published: (2024)
by: Zeng, Qingyuan, et al.
Published: (2024)
Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization
by: Hu, Yupeng, et al.
Published: (2023)
by: Hu, Yupeng, et al.
Published: (2023)
Solution for Point Tracking Task of ECCV 2nd Perception Test Challenge 2024
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
by: Peng, Yingzhe, et al.
Published: (2024)
by: Peng, Yingzhe, et al.
Published: (2024)
Solution for OOD-CV UNICORN Challenge 2024 Object Detection Assistance LLM Counting Ability Improvement
by: Chi, Zhouyang, et al.
Published: (2024)
by: Chi, Zhouyang, et al.
Published: (2024)
Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
1st Place Solution of Multiview Egocentric Hand Tracking Challenge ECCV2024
by: Zou, Minqiang, et al.
Published: (2024)
by: Zou, Minqiang, et al.
Published: (2024)
Silver medal Solution for Image Matching Challenge 2024
by: Wang, Yian
Published: (2024)
by: Wang, Yian
Published: (2024)
Addressing Vulnerabilities in AI-Image Detection: Challenges and Proposed Solutions
by: Jiang, Justin
Published: (2024)
by: Jiang, Justin
Published: (2024)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
by: Chao, Dian, et al.
Published: (2024)
by: Chao, Dian, et al.
Published: (2024)
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic
by: Tang, Jianwei, et al.
Published: (2025)
by: Tang, Jianwei, et al.
Published: (2025)
VACT: A Video Automatic Causal Testing System and a Benchmark
by: Yang, Haotong, et al.
Published: (2025)
by: Yang, Haotong, et al.
Published: (2025)
Solution for CVPR 2024 UG2+ Challenge Track on All Weather Semantic Segmentation
by: Yu, Jun, et al.
Published: (2024)
by: Yu, Jun, et al.
Published: (2024)
Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution
by: Wei, Timothy, et al.
Published: (2024)
by: Wei, Timothy, et al.
Published: (2024)
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
by: Si, Shengyu, et al.
Published: (2026)
by: Si, Shengyu, et al.
Published: (2026)
Equal is Not Always Fair: A New Perspective on Hyperspectral Representation Non-Uniformity
by: Quan, Wuzhou, et al.
Published: (2025)
by: Quan, Wuzhou, et al.
Published: (2025)
Mitigating Query Selection Bias in Referring Video Object Segmentation
by: Zhang, Dingwei, et al.
Published: (2025)
by: Zhang, Dingwei, et al.
Published: (2025)
Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
by: Larchenko, Ilia, et al.
Published: (2025)
by: Larchenko, Ilia, et al.
Published: (2025)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
by: Yang, Siyuan, et al.
Published: (2023)
by: Yang, Siyuan, et al.
Published: (2023)
Explore the Hallucination on Low-level Perception for MLLMs
by: Sun, Yinan, et al.
Published: (2024)
by: Sun, Yinan, et al.
Published: (2024)
3rd Place Solution for MOSE Track in CVPR 2024 PVUW workshop: Complex Video Object Segmentation
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
by: Cheng, Jen-Hao, et al.
Published: (2025)
by: Cheng, Jen-Hao, et al.
Published: (2025)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
MALT: Multi-scale Action Learning Transformer for Online Action Detection
by: Yang, Zhipeng, et al.
Published: (2024)
by: Yang, Zhipeng, et al.
Published: (2024)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
Human-AI Collaboration Mechanism Study on AIGC Assisted Image Production for Special Coverage
by: Yang, Yajie, et al.
Published: (2025)
by: Yang, Yajie, et al.
Published: (2025)
A Decade of Action Quality Assessment: Largest Systematic Survey of Trends, Challenges, and Future Directions
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
PEnG: Pose-Enhanced Geo-Localisation
by: Shore, Tavis, et al.
Published: (2024)
by: Shore, Tavis, et al.
Published: (2024)
ME-CPT: Multi-Task Enhanced Cross-Temporal Point Transformer for Urban 3D Change Detection
by: Zhang, Luqi, et al.
Published: (2025)
by: Zhang, Luqi, et al.
Published: (2025)
MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
DDaTR: Dynamic Difference-aware Temporal Residual Network for Longitudinal Radiology Report Generation
by: Song, Shanshan, et al.
Published: (2025)
by: Song, Shanshan, et al.
Published: (2025)
Similar Items
-
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
by: Huang, Yurui, et al.
Published: (2024) -
The Solution for Single Object Tracking Task of Perception Test Challenge 2024
by: Zhong, Zhiqiang, et al.
Published: (2024) -
Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024
by: Ahn, Jinwoo, et al.
Published: (2024) -
Cross-Task Attack: A Self-Supervision Generative Framework Based on Attention Shift
by: Zeng, Qingyuan, et al.
Published: (2024) -
Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization
by: Hu, Yupeng, et al.
Published: (2023)