GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Rui, Zhu, Lianghui, Zhang, Yuxuan, Cheng, Tianheng, Liu, Lei, Liu, Heng, Ran, Longjin, Chen, Xiaoxin, Liu, Wenyu, Wang, Xinggang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
by: Zhu, Lianghui, et al.
Published: (2025)
by: Zhu, Lianghui, et al.
Published: (2025)
TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
by: Li, Zongming, et al.
Published: (2025)
by: Li, Zongming, et al.
Published: (2025)
ControlAR: Controllable Image Generation with Autoregressive Models
by: Li, Zongming, et al.
Published: (2024)
by: Li, Zongming, et al.
Published: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
by: Li, Yongkang, et al.
Published: (2024)
by: Li, Yongkang, et al.
Published: (2024)
GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
by: Xu, Ziyang, et al.
Published: (2024)
by: Xu, Ziyang, et al.
Published: (2024)
Occupancy as Set of Points
by: Shi, Yiang, et al.
Published: (2024)
by: Shi, Yiang, et al.
Published: (2024)
MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
by: Zhang, Wenrui, et al.
Published: (2025)
by: Zhang, Wenrui, et al.
Published: (2025)
PixelHacker: Image Inpainting with Structural and Semantic Consistency
by: Xu, Ziyang, et al.
Published: (2025)
by: Xu, Ziyang, et al.
Published: (2025)
WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
Polar Parametrization for Vision-based Surround-View 3D Detection
by: Chen, Shaoyu, et al.
Published: (2022)
by: Chen, Shaoyu, et al.
Published: (2022)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
by: Zhu, Lianghui, et al.
Published: (2023)
by: Zhu, Lianghui, et al.
Published: (2023)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
GaitGS: Temporal Feature Learning in Granularity and Span Dimension for Gait Recognition
by: Xiong, Haijun, et al.
Published: (2023)
by: Xiong, Haijun, et al.
Published: (2023)
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
by: Jiang, Haoyi, et al.
Published: (2024)
by: Jiang, Haoyi, et al.
Published: (2024)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
by: Hu, Bin, et al.
Published: (2024)
by: Hu, Bin, et al.
Published: (2024)
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
by: Zong, Yongshuo, et al.
Published: (2025)
by: Zong, Yongshuo, et al.
Published: (2025)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
by: Wu, Xianfeng, et al.
Published: (2025)
by: Wu, Xianfeng, et al.
Published: (2025)
TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion Generation
by: Cao, Yiyang, et al.
Published: (2026)
by: Cao, Yiyang, et al.
Published: (2026)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
by: Liao, Bencheng, et al.
Published: (2025)
by: Liao, Bencheng, et al.
Published: (2025)
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
by: Yao, Jingfeng, et al.
Published: (2024)
by: Yao, Jingfeng, et al.
Published: (2024)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
by: Zou, Jialv, et al.
Published: (2024)
by: Zou, Jialv, et al.
Published: (2024)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
by: Li, Yingyue, et al.
Published: (2025)
by: Li, Yingyue, et al.
Published: (2025)
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
by: Yao, Jingfeng, et al.
Published: (2023)
by: Yao, Jingfeng, et al.
Published: (2023)
Causality-inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
by: Xiong, Haijun, et al.
Published: (2024)
by: Xiong, Haijun, et al.
Published: (2024)
DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
by: Zhu, Yueting, et al.
Published: (2025)
by: Zhu, Yueting, et al.
Published: (2025)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding
by: Hui, Tianrui, et al.
Published: (2023)
by: Hui, Tianrui, et al.
Published: (2023)
DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
by: Kasuba, Badri Vishal, et al.
Published: (2025)
by: Kasuba, Badri Vishal, et al.
Published: (2025)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
by: Kim, Jeonghwan, et al.
Published: (2026)
by: Kim, Jeonghwan, et al.
Published: (2026)
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
by: Su, Yuejiao, et al.
Published: (2025)
by: Su, Yuejiao, et al.
Published: (2025)
MoSt-DSA: Modeling Motion and Structural Interactions for Direct Multi-Frame Interpolation in DSA Images
by: Xu, Ziyang, et al.
Published: (2024)
by: Xu, Ziyang, et al.
Published: (2024)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
by: Shabbir, Akashah, et al.
Published: (2025)
by: Shabbir, Akashah, et al.
Published: (2025)
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
by: Shu, Yan, et al.
Published: (2026)
by: Shu, Yan, et al.
Published: (2026)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
Similar Items
-
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024) -
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
by: Zhu, Lianghui, et al.
Published: (2025) -
TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
by: Li, Zongming, et al.
Published: (2025) -
ControlAR: Controllable Image Generation with Autoregressive Models
by: Li, Zongming, et al.
Published: (2024) -
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)