SegLLM: Multi-round Reasoning Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, XuDong, Zhang, Shaolun, Li, Shufan, Kallidromitis, Konstantinos, Li, Kehan, Kato, Yusuke, Kozuka, Kazuki, Darrell, Trevor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Aligning Diffusion Models by Optimizing Human Utility
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization
by: Li, Shufan, et al.
Published: (2026)
by: Li, Shufan, et al.
Published: (2026)
Hyperbolic Active Learning for Semantic Segmentation under Domain Shift
by: Franco, Luca, et al.
Published: (2023)
by: Franco, Luca, et al.
Published: (2023)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
MobileWorldBench: Towards Semantic World Modeling For Mobile Agents
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
xT: Nested Tokenization for Larger Context in Large Images
by: Gupta, Ritwik, et al.
Published: (2024)
by: Gupta, Ritwik, et al.
Published: (2024)
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation
by: Wen, Junwei, et al.
Published: (2026)
by: Wen, Junwei, et al.
Published: (2026)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
VideoSeg-R1:Reasoning Video Object Segmentation via Reinforcement Learning
by: Xu, Zishan, et al.
Published: (2025)
by: Xu, Zishan, et al.
Published: (2025)
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation
by: Qian, Rui, et al.
Published: (2026)
by: Qian, Rui, et al.
Published: (2026)
LLM-Seg: Bridging Image Segmentation and Large Language Model Reasoning
by: Wang, Junchi, et al.
Published: (2024)
by: Wang, Junchi, et al.
Published: (2024)
Visually Prompted Benchmarks Are Surprisingly Fragile
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Constantly Improving Image Models Need Constantly Improving Benchmarks
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
SteerSeg: Attention Steering for Reasoning Video Segmentation
by: Cheraghian, Ali, et al.
Published: (2026)
by: Cheraghian, Ali, et al.
Published: (2026)
Latent Implicit Visual Reasoning
by: Li, Kelvin, et al.
Published: (2025)
by: Li, Kelvin, et al.
Published: (2025)
MarsSeg: Mars Surface Semantic Segmentation with Multi-level Extractor and Connector
by: Li, Junbo, et al.
Published: (2024)
by: Li, Junbo, et al.
Published: (2024)
MedSeg-R: Medical Image Segmentation with Clinical Reasoning
by: Shao, Hao, et al.
Published: (2025)
by: Shao, Hao, et al.
Published: (2025)
Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search
by: Liang, Tianming, et al.
Published: (2026)
by: Liang, Tianming, et al.
Published: (2026)
FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
by: Zhang, Zekang, et al.
Published: (2026)
by: Zhang, Zekang, et al.
Published: (2026)
CoT-Seg: Rethinking Segmentation with Chain-of-Thought Reasoning and Self-Correction
by: Kao, Shiu-hong, et al.
Published: (2026)
by: Kao, Shiu-hong, et al.
Published: (2026)
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
by: Qu, Jinyuan, et al.
Published: (2026)
by: Qu, Jinyuan, et al.
Published: (2026)
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
by: Jisheng, Dang, et al.
Published: (2025)
by: Jisheng, Dang, et al.
Published: (2025)
Wild2Avatar: Rendering Humans Behind Occlusions
by: Xiang, Tiange, et al.
Published: (2023)
by: Xiang, Tiange, et al.
Published: (2023)
OMG-Seg: Is One Model Good Enough For All Segmentation?
by: Li, Xiangtai, et al.
Published: (2024)
by: Li, Xiangtai, et al.
Published: (2024)
Pose2Seg: Detection Free Human Instance Segmentation
by: Zhang, Song-Hai, et al.
Published: (2018)
by: Zhang, Song-Hai, et al.
Published: (2018)
WOW-Seg: A Word-free Open World Segmentation Model
by: Li, Danyang, et al.
Published: (2026)
by: Li, Danyang, et al.
Published: (2026)
Mamba-ND: Selective State Space Modeling for Multi-Dimensional Data
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
Two-stage Rainfall-Forecasting Diffusion Model
by: Ling, XuDong, et al.
Published: (2024)
by: Ling, XuDong, et al.
Published: (2024)
EyeSeg: An Uncertainty-Aware Eye Segmentation Framework for AR/VR
by: Peng, Zhengyuan, et al.
Published: (2025)
by: Peng, Zhengyuan, et al.
Published: (2025)
ACS-SegNet: An Attention-Based CNN-SegFormer Segmentation Network for Tissue Segmentation in Histopathology
by: Torbati, Nima, et al.
Published: (2025)
by: Torbati, Nima, et al.
Published: (2025)
LaneSegNet: Map Learning with Lane Segment Perception for Autonomous Driving
by: Li, Tianyu, et al.
Published: (2023)
by: Li, Tianyu, et al.
Published: (2023)
DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models
by: He, Yulin, et al.
Published: (2026)
by: He, Yulin, et al.
Published: (2026)
MetaSeg: Content-Aware Meta-Net for Omni-Supervised Semantic Segmentation
by: Jiang, Shenwang, et al.
Published: (2024)
by: Jiang, Shenwang, et al.
Published: (2024)
Similar Items
-
Aligning Diffusion Models by Optimizing Human Utility
by: Li, Shufan, et al.
Published: (2024) -
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
by: Li, Shufan, et al.
Published: (2025) -
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024) -
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024) -
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
by: Li, Shufan, et al.
Published: (2025)