Reasoning to Attend: Try to Understand How <SEG> Token Works
Fuente:
arXiv
Saved in:
| Main Authors: | Qian, Rui, Yin, Xin, Dou, Dejing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UGround: Towards Unified Visual Grounding with Unrolled Transformers
by: Qian, Rui, et al.
Published: (2025)
by: Qian, Rui, et al.
Published: (2025)
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation
by: Qian, Rui, et al.
Published: (2026)
by: Qian, Rui, et al.
Published: (2026)
Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
by: Fang, Haipeng, et al.
Published: (2025)
by: Fang, Haipeng, et al.
Published: (2025)
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
by: Shi, Baifeng, et al.
Published: (2026)
by: Shi, Baifeng, et al.
Published: (2026)
SEG-SAM: Semantic-Guided SAM for Unified Medical Image Segmentation
by: Huang, Shuangping, et al.
Published: (2024)
by: Huang, Shuangping, et al.
Published: (2024)
CrossVTON: Mimicking the Logic Reasoning on Cross-category Virtual Try-on guided by Tri-zone Priors
by: Luo, Donghao, et al.
Published: (2025)
by: Luo, Donghao, et al.
Published: (2025)
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
by: Chen, Haodong, et al.
Published: (2025)
by: Chen, Haodong, et al.
Published: (2025)
Locality-Attending Vision Transformer
by: Hajimiri, Sina, et al.
Published: (2026)
by: Hajimiri, Sina, et al.
Published: (2026)
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
by: Zhang, Renshan, et al.
Published: (2024)
by: Zhang, Renshan, et al.
Published: (2024)
TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
by: Xing, Jiazheng, et al.
Published: (2024)
by: Xing, Jiazheng, et al.
Published: (2024)
RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
by: Wang, Zixun, et al.
Published: (2025)
by: Wang, Zixun, et al.
Published: (2025)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
by: Jiang, Xixi, et al.
Published: (2025)
by: Jiang, Xixi, et al.
Published: (2025)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
EarlyTom: Early Token Compression Completes Fast Video Understanding
by: Wang, Hesong, et al.
Published: (2026)
by: Wang, Hesong, et al.
Published: (2026)
TAMM: TriAdapter Multi-Modal Learning for 3D Shape Understanding
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
by: Liu, Wutao, et al.
Published: (2025)
by: Liu, Wutao, et al.
Published: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
by: Jiang, Jindong, et al.
Published: (2025)
by: Jiang, Jindong, et al.
Published: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
by: Qi, Haozhe, et al.
Published: (2026)
by: Qi, Haozhe, et al.
Published: (2026)
GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning
by: Miao, Deshui, et al.
Published: (2026)
by: Miao, Deshui, et al.
Published: (2026)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
by: Mao, Weian, et al.
Published: (2026)
by: Mao, Weian, et al.
Published: (2026)
OOD-SEG: Exploiting out-of-distribution detection techniques for learning image segmentation from sparse multi-class positive-only annotations
by: Wang, Junwen, et al.
Published: (2024)
by: Wang, Junwen, et al.
Published: (2024)
Shape-Guided Clothing Warping for Virtual Try-On
by: Han, Xiaoyu, et al.
Published: (2025)
by: Han, Xiaoyu, et al.
Published: (2025)
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
by: Guo, Hanzhong, et al.
Published: (2024)
by: Guo, Hanzhong, et al.
Published: (2024)
OmniTry: Virtual Try-On Anything without Masks
by: Feng, Yutong, et al.
Published: (2025)
by: Feng, Yutong, et al.
Published: (2025)
BC-MRI-SEG: A Breast Cancer MRI Tumor Segmentation Benchmark
by: Bilic, Anthony, et al.
Published: (2024)
by: Bilic, Anthony, et al.
Published: (2024)
Teeth-SEG: An Efficient Instance Segmentation Framework for Orthodontic Treatment based on Anthropic Prior Knowledge
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
by: Pan, Liang, et al.
Published: (2025)
by: Pan, Liang, et al.
Published: (2025)
Principles of Visual Tokens for Efficient Video Understanding
by: Hao, Xinyue, et al.
Published: (2024)
by: Hao, Xinyue, et al.
Published: (2024)
Towards the Automatic Segmentation, Modeling and Meshing of the Aortic Vessel Tree from Multicenter Acquisitions: An Overview of the SEG.A. 2023 Segmentation of the Aorta Challenge
by: Jin, Yuan, et al.
Published: (2025)
by: Jin, Yuan, et al.
Published: (2025)
Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening
by: Li, Junfeng, et al.
Published: (2026)
by: Li, Junfeng, et al.
Published: (2026)
When Token Pruning is Worse than Random: Understanding Visual Token Information in VLLMs
by: Wang, Yahong, et al.
Published: (2025)
by: Wang, Yahong, et al.
Published: (2025)
2D Gaussians Meet Visual Tokenizer
by: Shi, Yiang, et al.
Published: (2025)
by: Shi, Yiang, et al.
Published: (2025)
SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
by: Zheng, Hongpei, et al.
Published: (2025)
by: Zheng, Hongpei, et al.
Published: (2025)
GlamTry: Advancing Virtual Try-On for High-End Accessories
by: Chang, Ting-Yu, et al.
Published: (2024)
by: Chang, Ting-Yu, et al.
Published: (2024)
Tri$^{2}$-plane: Thinking Head Avatar via Feature Pyramid
by: Song, Luchuan, et al.
Published: (2024)
by: Song, Luchuan, et al.
Published: (2024)
Fast SceneScript: Fast and Accurate Language-Based 3D Scene Understanding via Multi-Token Prediction
by: Yin, Ruihong, et al.
Published: (2025)
by: Yin, Ruihong, et al.
Published: (2025)
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
by: Liu, Man, et al.
Published: (2024)
by: Liu, Man, et al.
Published: (2024)
Video, How Do Your Tokens Merge?
by: Pollard, Sam, et al.
Published: (2025)
by: Pollard, Sam, et al.
Published: (2025)
Attend to what I say: Highlighting relevant content on slides
by: M, Megha Mariam K, et al.
Published: (2026)
by: M, Megha Mariam K, et al.
Published: (2026)
Similar Items
-
UGround: Towards Unified Visual Grounding with Unrolled Transformers
by: Qian, Rui, et al.
Published: (2025) -
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation
by: Qian, Rui, et al.
Published: (2026) -
Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
by: Fang, Haipeng, et al.
Published: (2025) -
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025) -
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
by: Shi, Baifeng, et al.
Published: (2026)