TAPTR: Tracking Any Point with Transformers as Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongyang, Zhang, Hao, Liu, Shilong, Zeng, Zhaoyang, Ren, Tianhe, Li, Feng, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TAPTRv2: Attention-based Position Update Improves Tracking Any Point
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024)
by: Qu, Jinyuan, et al.
Published: (2024)
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
ETAP: Event-based Tracking of Any Point
by: Hamann, Friedhelm, et al.
Published: (2024)
by: Hamann, Friedhelm, et al.
Published: (2024)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
by: Fan, Xianzhe, et al.
Published: (2026)
by: Fan, Xianzhe, et al.
Published: (2026)
Detect Anything via Next Point Prediction
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer
by: Ding, Tianye, et al.
Published: (2024)
by: Ding, Tianye, et al.
Published: (2024)
SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
by: Qu, Jinyuan, et al.
Published: (2025)
by: Qu, Jinyuan, et al.
Published: (2025)
Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2024)
by: Zhang, Guowen, et al.
Published: (2024)
The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Cross-Level Sensor Fusion with Object Lists via Transformer for 3D Object Detection
by: Liu, Xiangzhong, et al.
Published: (2025)
by: Liu, Xiangzhong, et al.
Published: (2025)
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
by: Pan, Yue, et al.
Published: (2025)
by: Pan, Yue, et al.
Published: (2025)
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Point Tree Transformer for Point Cloud Registration
by: Wang, Meiling, et al.
Published: (2024)
by: Wang, Meiling, et al.
Published: (2024)
TrackVLA: Embodied Visual Tracking in the Wild
by: Wang, Shaoan, et al.
Published: (2025)
by: Wang, Shaoan, et al.
Published: (2025)
AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond
by: Zhang, Haiming, et al.
Published: (2026)
by: Zhang, Haiming, et al.
Published: (2026)
Griffin: Aerial-Ground Cooperative Detection and Tracking Dataset and Benchmark
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
FastMAC: Stochastic Spectral Sampling of Correspondence Graph
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
MAS-SAM: Segment Any Marine Animal with Aggregated Features
by: Yan, Tianyu, et al.
Published: (2024)
by: Yan, Tianyu, et al.
Published: (2024)
Kalib: Easy Hand-Eye Calibration with Reference Point Tracking
by: Tang, Tutian, et al.
Published: (2024)
by: Tang, Tutian, et al.
Published: (2024)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera
by: Guo, Yuliang, et al.
Published: (2025)
by: Guo, Yuliang, et al.
Published: (2025)
Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation
by: Zhang, Hai, et al.
Published: (2026)
by: Zhang, Hai, et al.
Published: (2026)
GarmentTracking: Category-Level Garment Pose Tracking
by: Xue, Han, et al.
Published: (2023)
by: Xue, Han, et al.
Published: (2023)
Leveraging Previous-Traversal Point Cloud Map Priors for Camera-Based 3D Object Detection and Tracking
by: Käppeler, Markus, et al.
Published: (2026)
by: Käppeler, Markus, et al.
Published: (2026)
Weakly Supervised Point Clouds Transformer for 3D Object Detection
by: Tang, Zuojin, et al.
Published: (2023)
by: Tang, Zuojin, et al.
Published: (2023)
Loop Closure using AnyLoc Visual Place Recognition in DPV-SLAM
by: Zhang, Wenzheng, et al.
Published: (2026)
by: Zhang, Wenzheng, et al.
Published: (2026)
Multi-Modal UAV Detection, Classification and Tracking Algorithm -- Technical Report for CVPR 2024 UG2 Challenge
by: Deng, Tianchen, et al.
Published: (2024)
by: Deng, Tianchen, et al.
Published: (2024)
Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation
by: Bharadhwaj, Homanga, et al.
Published: (2024)
by: Bharadhwaj, Homanga, et al.
Published: (2024)
Fast Point Cloud to Mesh Reconstruction for Deformable Object Tracking
by: Mansour, Elham Amin, et al.
Published: (2023)
by: Mansour, Elham Amin, et al.
Published: (2023)
AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation
by: Tan, Hengkai, et al.
Published: (2025)
by: Tan, Hengkai, et al.
Published: (2025)
AnyImageNav: Any-View Geometry for Precise Last-Meter Image-Goal Navigation
by: Deng, Yijie, et al.
Published: (2026)
by: Deng, Yijie, et al.
Published: (2026)
DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos
by: Li, Can, et al.
Published: (2026)
by: Li, Can, et al.
Published: (2026)
Interpretable Affordance Detection on 3D Point Clouds with Probabilistic Prototypes
by: Li, Maximilian Xiling, et al.
Published: (2025)
by: Li, Maximilian Xiling, et al.
Published: (2025)
SpatialPoint: Spatial-aware Point Prediction for Embodied Localization
by: Zhu, Qiming, et al.
Published: (2026)
by: Zhu, Qiming, et al.
Published: (2026)
Tracking Tumors under Deformation from Partial Point Clouds using Occupancy Networks
by: Henrich, Pit, et al.
Published: (2024)
by: Henrich, Pit, et al.
Published: (2024)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
by: Cui, Jieming, et al.
Published: (2024)
by: Cui, Jieming, et al.
Published: (2024)
RockTrack: A 3D Robust Multi-Camera-Ken Multi-Object Tracking Framework
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
Similar Items
-
TAPTRv2: Attention-based Position Update Improves Tracking Any Point
by: Li, Hongyang, et al.
Published: (2024) -
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024) -
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
by: Jiang, Qing, et al.
Published: (2024) -
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025) -
ETAP: Event-based Tracking of Any Point
by: Hamann, Friedhelm, et al.
Published: (2024)