Dense Optical Tracking: Connecting the Dots
Fuente:
arXiv
Salvato in:
| Autori principali: | Moing, Guillaume Le, Ponce, Jean, Schmid, Cordelia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Online 3D Scene Reconstruction Using Neural Object Priors
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
di: Chabal, Thomas, et al.
Pubblicazione: (2025)
Dense Video Object Captioning from Disjoint Supervision
di: Zhou, Xingyi, et al.
Pubblicazione: (2023)
di: Zhou, Xingyi, et al.
Pubblicazione: (2023)
BrickNet: Graph-Backed Generative Brick Assembly
di: Kulits, Peter, et al.
Pubblicazione: (2026)
di: Kulits, Peter, et al.
Pubblicazione: (2026)
Streaming Dense Video Captioning
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
di: Fiastre, Gabriel, et al.
Pubblicazione: (2025)
di: Fiastre, Gabriel, et al.
Pubblicazione: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
Learning text-to-video retrieval from image captioning
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
di: Khan, Zeeshan, et al.
Pubblicazione: (2025)
di: Khan, Zeeshan, et al.
Pubblicazione: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2025)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
di: Garcia, Ricardo, et al.
Pubblicazione: (2024)
di: Garcia, Ricardo, et al.
Pubblicazione: (2024)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
Retrieval-Enhanced Contrastive Vision-Text Models
di: Iscen, Ahmet, et al.
Pubblicazione: (2023)
di: Iscen, Ahmet, et al.
Pubblicazione: (2023)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
di: Chen, Shizhe, et al.
Pubblicazione: (2024)
di: Chen, Shizhe, et al.
Pubblicazione: (2024)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
di: Wysoczańska, Monika, et al.
Pubblicazione: (2025)
di: Wysoczańska, Monika, et al.
Pubblicazione: (2025)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
Dense Matchers for Dense Tracking
di: Jelínek, Tomáš, et al.
Pubblicazione: (2024)
di: Jelínek, Tomáš, et al.
Pubblicazione: (2024)
HORT: Monocular Hand-held Objects Reconstruction with Transformers
di: Chen, Zerui, et al.
Pubblicazione: (2025)
di: Chen, Zerui, et al.
Pubblicazione: (2025)
Learning Correlation Structures for Vision Transformers
di: Kim, Manjin, et al.
Pubblicazione: (2024)
di: Kim, Manjin, et al.
Pubblicazione: (2024)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
di: Pacaud, Paul, et al.
Pubblicazione: (2025)
di: Pacaud, Paul, et al.
Pubblicazione: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
di: Ghosh, Partha, et al.
Pubblicazione: (2024)
di: Ghosh, Partha, et al.
Pubblicazione: (2024)
What Are You Doing? A Closer Look at Controllable Human Video Generation
di: Bugliarello, Emanuele, et al.
Pubblicazione: (2025)
di: Bugliarello, Emanuele, et al.
Pubblicazione: (2025)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
di: Kang, Dahyun, et al.
Pubblicazione: (2025)
di: Kang, Dahyun, et al.
Pubblicazione: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
di: Chen, Shizhe, et al.
Pubblicazione: (2025)
di: Chen, Shizhe, et al.
Pubblicazione: (2025)
MetricNet: Recovering Metric Scale in Generative Navigation Policies
di: Nayak, Abhijeet, et al.
Pubblicazione: (2025)
di: Nayak, Abhijeet, et al.
Pubblicazione: (2025)
Online Dense Point Tracking with Streaming Memory
di: Dong, Qiaole, et al.
Pubblicazione: (2025)
di: Dong, Qiaole, et al.
Pubblicazione: (2025)
SAMIDARE: Advanced Tracking-by-Segmentation for Dense Scenarios
di: Hirano, Shozaburo, et al.
Pubblicazione: (2026)
di: Hirano, Shozaburo, et al.
Pubblicazione: (2026)
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
di: Huang, Ian, et al.
Pubblicazione: (2025)
di: Huang, Ian, et al.
Pubblicazione: (2025)
DELTAv2: Accelerating Dense 3D Tracking
di: Ngo, Tuan Duc, et al.
Pubblicazione: (2025)
di: Ngo, Tuan Duc, et al.
Pubblicazione: (2025)
Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
di: Huang, Jiaqi, et al.
Pubblicazione: (2025)
DataDream: Few-shot Guided Dataset Generation
di: Kim, Jae Myung, et al.
Pubblicazione: (2024)
di: Kim, Jae Myung, et al.
Pubblicazione: (2024)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
di: Chen, Zerui, et al.
Pubblicazione: (2026)
di: Chen, Zerui, et al.
Pubblicazione: (2026)
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
di: Nam, Jisu, et al.
Pubblicazione: (2026)
di: Nam, Jisu, et al.
Pubblicazione: (2026)
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
di: Lu, Jiahao, et al.
Pubblicazione: (2026)
di: Lu, Jiahao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Online 3D Scene Reconstruction Using Neural Object Priors
di: Chabal, Thomas, et al.
Pubblicazione: (2025) -
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
di: Chabal, Thomas, et al.
Pubblicazione: (2025) -
Dense Video Object Captioning from Disjoint Supervision
di: Zhou, Xingyi, et al.
Pubblicazione: (2023) -
BrickNet: Graph-Backed Generative Brick Assembly
di: Kulits, Peter, et al.
Pubblicazione: (2026) -
Streaming Dense Video Captioning
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)