Detect Anything via Next Point Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Qing, Huo, Junan, Chen, Xingyu, Xiong, Yuda, Zeng, Zhaoyang, Chen, Yihao, Ren, Tianhe, Yu, Junzhi, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
TAPTR: Tracking Any Point with Transformers as Detection
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024)
by: Qu, Jinyuan, et al.
Published: (2024)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
TAPTRv2: Attention-based Position Update Improves Tracking Any Point
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
HandOS: 3D Hand Reconstruction in One Stage
by: Chen, Xingyu, et al.
Published: (2024)
by: Chen, Xingyu, et al.
Published: (2024)
LET-US: Long Event-Text Understanding of Scenes
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
by: Qu, Jinyuan, et al.
Published: (2025)
by: Qu, Jinyuan, et al.
Published: (2025)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
NuNext: Reframing Nucleus Detection as Next-Point Detection
by: Shui, Zhongyi, et al.
Published: (2026)
by: Shui, Zhongyi, et al.
Published: (2026)
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality Assumption
by: Chen, Du, et al.
Published: (2025)
by: Chen, Du, et al.
Published: (2025)
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
by: Shu, Han, et al.
Published: (2023)
by: Shu, Han, et al.
Published: (2023)
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
by: Chen, Zhuohao, et al.
Published: (2026)
by: Chen, Zhuohao, et al.
Published: (2026)
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
SNAP: Towards Segmenting Anything in Any Point Cloud
by: Gupta, Aniket, et al.
Published: (2025)
by: Gupta, Aniket, et al.
Published: (2025)
GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation
by: Lan, Yushi, et al.
Published: (2024)
by: Lan, Yushi, et al.
Published: (2024)
Multimodal OCR: Parse Anything from Documents
by: Zheng, Handong, et al.
Published: (2026)
by: Zheng, Handong, et al.
Published: (2026)
Autoregressive Video Generation beyond Next Frames Prediction
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
SAM3-UNet: Simplified Adaptation of Segment Anything Model 3
by: Xiong, Xinyu, et al.
Published: (2025)
by: Xiong, Xinyu, et al.
Published: (2025)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
by: Li, Aiden Yiliu, et al.
Published: (2025)
by: Li, Aiden Yiliu, et al.
Published: (2025)
Detect Anything 3D in the Wild
by: Zhang, Hanxue, et al.
Published: (2025)
by: Zhang, Hanxue, et al.
Published: (2025)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
by: Wu, Huimin, et al.
Published: (2025)
by: Wu, Huimin, et al.
Published: (2025)
SmartEraser: Remove Anything from Images using Masked-Region Guidance
by: Jiang, Longtao, et al.
Published: (2025)
by: Jiang, Longtao, et al.
Published: (2025)
Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework
by: Yao, Yumeng, et al.
Published: (2026)
by: Yao, Yumeng, et al.
Published: (2026)
Consistent-Point: Consistent Pseudo-Points for Semi-Supervised Crowd Counting and Localization
by: Zou, Yuda, et al.
Published: (2025)
by: Zou, Yuda, et al.
Published: (2025)
Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction
by: Chang, Chun-Peng, et al.
Published: (2026)
by: Chang, Chun-Peng, et al.
Published: (2026)
Enhancing the Reliability of Segment Anything Model for Auto-Prompting Medical Image Segmentation with Uncertainty Rectification
by: Zhang, Yichi, et al.
Published: (2023)
by: Zhang, Yichi, et al.
Published: (2023)
Segment Anything, Even Occluded
by: Tai, Wei-En, et al.
Published: (2025)
by: Tai, Wei-En, et al.
Published: (2025)
VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
by: Wu, Tianhe, et al.
Published: (2025)
by: Wu, Tianhe, et al.
Published: (2025)
Arbitrary Ratio Feature Compression via Next Token Prediction
by: Liu, Yufan, et al.
Published: (2026)
by: Liu, Yufan, et al.
Published: (2026)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
TAPNext: Tracking Any Point (TAP) as Next Token Prediction
by: Zholus, Artem, et al.
Published: (2025)
by: Zholus, Artem, et al.
Published: (2025)
Composition Vision-Language Understanding via Segment and Depth Anything Model
by: Huo, Mingxiao, et al.
Published: (2024)
by: Huo, Mingxiao, et al.
Published: (2024)
Similar Items
-
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
by: Jiang, Qing, et al.
Published: (2024) -
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025) -
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
by: Jiang, Qing, et al.
Published: (2025) -
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
by: Jiang, Qing, et al.
Published: (2024) -
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)