ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Yifan, Zhao, Zeyang, Gong, Yihong, Wei, Xing |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation
by: Zhao, Zeyang, et al.
Published: (2024)
by: Zhao, Zeyang, et al.
Published: (2024)
FARTrack: Fast Autoregressive Visual Tracking with High Performance
by: Wang, Guijie, et al.
Published: (2026)
by: Wang, Guijie, et al.
Published: (2026)
IOTA: Corrective Knowledge-Guided Prompt Learning via Black-White Box Framework
by: Wang, Shaokun, et al.
Published: (2026)
by: Wang, Shaokun, et al.
Published: (2026)
SpatialTrackerV2: 3D Point Tracking Made Easy
by: Xiao, Yuxi, et al.
Published: (2025)
by: Xiao, Yuxi, et al.
Published: (2025)
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)
by: Fuller, Anthony, et al.
Published: (2025)
Pay Attention to Where You Looked
by: Berian, Alex, et al.
Published: (2026)
by: Berian, Alex, et al.
Published: (2026)
Curriculum Dataset Distillation
by: Ma, Zhiheng, et al.
Published: (2024)
by: Ma, Zhiheng, et al.
Published: (2024)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Beyond Prompt Learning: Continual Adapter for Efficient Rehearsal-Free Continual Learning
by: Gao, Xinyuan, et al.
Published: (2024)
by: Gao, Xinyuan, et al.
Published: (2024)
PhyTracker: An Online Tracker for Phytoplankton
by: Yu, Yang, et al.
Published: (2024)
by: Yu, Yang, et al.
Published: (2024)
Few-shot Online Anomaly Detection and Segmentation
by: Wei, Shenxing, et al.
Published: (2024)
by: Wei, Shenxing, et al.
Published: (2024)
Where do Large Vision-Language Models Look at when Answering Questions?
by: Xing, Xiaoying, et al.
Published: (2025)
by: Xing, Xiaoying, et al.
Published: (2025)
Autoregressive Image Generation with Vision Full-view Prompt
by: Cai, Miaomiao, et al.
Published: (2025)
by: Cai, Miaomiao, et al.
Published: (2025)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
Prompt-Agnostic Adversarial Perturbation for Customized Diffusion Models
by: Wan, Cong, et al.
Published: (2024)
by: Wan, Cong, et al.
Published: (2024)
Diversity Covariance-Aware Prompt Learning for Vision-Language Models
by: Dong, Songlin, et al.
Published: (2025)
by: Dong, Songlin, et al.
Published: (2025)
Optimizing Multi-Modality Trackers via Significance-Regularized Tuning
by: Chen, Zhiwen, et al.
Published: (2025)
by: Chen, Zhiwen, et al.
Published: (2025)
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
by: Zhang, Zhengbo, et al.
Published: (2024)
by: Zhang, Zhengbo, et al.
Published: (2024)
TACOcc:Target-Adaptive Cross-Modal Fusion with Volume Rendering for 3D Semantic Occupancy
by: Lei, Luyao, et al.
Published: (2025)
by: Lei, Luyao, et al.
Published: (2025)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
Positional Prompt Tuning for Efficient 3D Representation Learning
by: Zhang, Shaochen, et al.
Published: (2024)
by: Zhang, Shaochen, et al.
Published: (2024)
iKUN: Speak to Trackers without Retraining
by: Du, Yunhao, et al.
Published: (2023)
by: Du, Yunhao, et al.
Published: (2023)
DescribeEarth: Describe Anything for Remote Sensing Images
by: Li, Kaiyu, et al.
Published: (2025)
by: Li, Kaiyu, et al.
Published: (2025)
Described Spatial-Temporal Video Detection
by: Ji, Wei, et al.
Published: (2024)
by: Ji, Wei, et al.
Published: (2024)
SAM2-SGP: Enhancing SAM2 for Medical Image Segmentation via Support-Set Guided Prompting
by: Xing, Yang, et al.
Published: (2025)
by: Xing, Yang, et al.
Published: (2025)
Benchmarking SAM2-based Trackers on FMOX
by: Aktas, Senem, et al.
Published: (2025)
by: Aktas, Senem, et al.
Published: (2025)
ReMoT: Reinforcement Learning with Motion Contrast Triplets
by: Wan, Cong, et al.
Published: (2026)
by: Wan, Cong, et al.
Published: (2026)
CEAT: Continual Expansion and Absorption Transformer for Non-Exemplar Class-Incremental Learning
by: Gao, Xinyuan, et al.
Published: (2024)
by: Gao, Xinyuan, et al.
Published: (2024)
Grid: Omni Visual Generation
by: Wan, Cong, et al.
Published: (2024)
by: Wan, Cong, et al.
Published: (2024)
Describe Anything in Medical Images
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
by: Duan, Yuxiang, et al.
Published: (2025)
by: Duan, Yuxiang, et al.
Published: (2025)
DecoFuse: Decomposing and Fusing the "What", "Where", and "How" for Brain-Inspired fMRI-to-Video Decoding
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
Multi-State Tracker: Enhancing Efficient Object Tracking via Multi-State Specialization and Interaction
by: Wang, Shilei, et al.
Published: (2025)
by: Wang, Shilei, et al.
Published: (2025)
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
by: Wang, Zhichuan, et al.
Published: (2025)
by: Wang, Zhichuan, et al.
Published: (2025)
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
by: Kothandaraman, Divya, et al.
Published: (2024)
by: Kothandaraman, Divya, et al.
Published: (2024)
DecoderTracker: Decoder-Only Method for Multiple-Object Tracking
by: Pan, Liao, et al.
Published: (2023)
by: Pan, Liao, et al.
Published: (2023)
Described Object Detection: Liberating Object Detection with Flexible Expressions
by: Xie, Chi, et al.
Published: (2023)
by: Xie, Chi, et al.
Published: (2023)
RAMCT: Novel Region-adaptive Multi-channel Tracker with Iterative Tikhonov Regularization for Thermal Infrared Tracking
by: Zhang, Shang, et al.
Published: (2025)
by: Zhang, Shang, et al.
Published: (2025)
MonoFormer: One Transformer for Both Diffusion and Autoregression
by: Zhao, Chuyang, et al.
Published: (2024)
by: Zhao, Chuyang, et al.
Published: (2024)
Similar Items
-
Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation
by: Zhao, Zeyang, et al.
Published: (2024) -
FARTrack: Fast Autoregressive Visual Tracking with High Performance
by: Wang, Guijie, et al.
Published: (2026) -
IOTA: Corrective Knowledge-Guided Prompt Learning via Black-White Box Framework
by: Wang, Shaokun, et al.
Published: (2026) -
SpatialTrackerV2: 3D Point Tracking Made Easy
by: Xiao, Yuxi, et al.
Published: (2025) -
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
by: Fuller, Anthony, et al.
Published: (2025)