TAPNext++: What's Next for Tracking Any Point (TAP)?
Fuente:
arXiv
Saved in:
| Main Authors: | Jung, Sebastian, Zholus, Artem, Sundermeyer, Martin, Doersch, Carl, Goroshin, Ross, Tan, David Joseph, Chandar, Sarath, Triebel, Rudolph, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
by: Zholus, Artem, et al.
Published: (2025)
by: Zholus, Artem, et al.
Published: (2025)
BootsTAP: Bootstrapped Training for Tracking-Any-Point
by: Doersch, Carl, et al.
Published: (2024)
by: Doersch, Carl, et al.
Published: (2024)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
Mastering Memory Tasks with World Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
MV-TAP: Tracking Any Point in Multi-View Videos
by: Koo, Jahyeok, et al.
Published: (2025)
by: Koo, Jahyeok, et al.
Published: (2025)
Towards Real-Time Open-Vocabulary Video Instance Segmentation
by: Yan, Bin, et al.
Published: (2024)
by: Yan, Bin, et al.
Published: (2024)
Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners
by: Araslanov, Nikita, et al.
Published: (2026)
by: Araslanov, Nikita, et al.
Published: (2026)
BindGPT: A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning
by: Zholus, Artem, et al.
Published: (2024)
by: Zholus, Artem, et al.
Published: (2024)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
by: Koppula, Skanda, et al.
Published: (2024)
by: Koppula, Skanda, et al.
Published: (2024)
Continuous Histogram Loss: Beyond Neural Similarity
by: Zholus, Artem, et al.
Published: (2020)
by: Zholus, Artem, et al.
Published: (2020)
Towards Explaining Uncertainty Estimates in Point Cloud Registration
by: Qin, Ziyuan, et al.
Published: (2024)
by: Qin, Ziyuan, et al.
Published: (2024)
The Expressive Limits of Diagonal SSMs for State-Tracking
by: Shakerinava, Mehran, et al.
Published: (2026)
by: Shakerinava, Mehran, et al.
Published: (2026)
AnthroTAP: Learning Point Tracking with Real-World Motion
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
NeoBERT: A Next-Generation BERT
by: Breton, Lola Le, et al.
Published: (2025)
by: Breton, Lola Le, et al.
Published: (2025)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
TRecViT: A Recurrent Video Transformer
by: Pătrăucean, Viorica, et al.
Published: (2024)
by: Pătrăucean, Viorica, et al.
Published: (2024)
One2Any: One-Reference 6D Pose Estimation for Any Object
by: Liu, Mengya, et al.
Published: (2025)
by: Liu, Mengya, et al.
Published: (2025)
Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image
by: Humt, Matthias, et al.
Published: (2025)
by: Humt, Matthias, et al.
Published: (2025)
A Learning-based Controller for Multi-Contact Grasps on Unknown Objects with a Dexterous Hand
by: Winkelbauer, Dominik, et al.
Published: (2024)
by: Winkelbauer, Dominik, et al.
Published: (2024)
RECALL: Rehearsal-free Continual Learning for Object Classification
by: Knauer, Markus, et al.
Published: (2022)
by: Knauer, Markus, et al.
Published: (2022)
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
by: Baldelli, Davide, et al.
Published: (2026)
by: Baldelli, Davide, et al.
Published: (2026)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
by: Guiroy, Simon, et al.
Published: (2025)
by: Guiroy, Simon, et al.
Published: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
Intelligent Switching for Reset-Free RL
by: Patil, Darshan, et al.
Published: (2024)
by: Patil, Darshan, et al.
Published: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Interpretability Needs a New Paradigm
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
by: Chitsaz, Kamran, et al.
Published: (2024)
by: Chitsaz, Kamran, et al.
Published: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
Lookbehind-SAM: k steps back, 1 step forward
by: Mordido, Gonçalo, et al.
Published: (2023)
by: Mordido, Gonçalo, et al.
Published: (2023)
RaNeuS: Ray-adaptive Neural Surface Reconstruction
by: Wang, Yida, et al.
Published: (2024)
by: Wang, Yida, et al.
Published: (2024)
Self-supervised Latent Space Optimization with Nebula Variational Coding
by: Wang, Yida, et al.
Published: (2025)
by: Wang, Yida, et al.
Published: (2025)
ETAP: Event-based Tracking of Any Point
by: Hamann, Friedhelm, et al.
Published: (2024)
by: Hamann, Friedhelm, et al.
Published: (2024)
TAPTR: Tracking Any Point with Transformers as Detection
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
Making the Flow Glow -- Robot Perception under Severe Lighting Conditions using Normalizing Flow Gradients
by: Lind, Simon Kristoffersson, et al.
Published: (2024)
by: Lind, Simon Kristoffersson, et al.
Published: (2024)
Human-Interpretable Uncertainty Explanations for Point Cloud Registration
by: Gaus, Johannes A., et al.
Published: (2025)
by: Gaus, Johannes A., et al.
Published: (2025)
Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
by: Liu, Jiaxiong, et al.
Published: (2024)
by: Liu, Jiaxiong, et al.
Published: (2024)
AnyUp: Universal Feature Upsampling
by: Wimmer, Thomas, et al.
Published: (2025)
by: Wimmer, Thomas, et al.
Published: (2025)
Similar Items
-
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
by: Zholus, Artem, et al.
Published: (2025) -
BootsTAP: Bootstrapped Training for Tracking-Any-Point
by: Doersch, Carl, et al.
Published: (2024) -
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
by: Nilaksh, et al.
Published: (2026) -
Mastering Memory Tasks with World Models
by: Samsami, Mohammad Reza, et al.
Published: (2024) -
MV-TAP: Tracking Any Point in Multi-View Videos
by: Koo, Jahyeok, et al.
Published: (2025)