CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Zhuoyan, Wu, Yinghao, Cheng, Tianheng, Liu, Yong, Xiao, Yicheng, Wang, Hongfa, Zhang, Xiao-Ping, Yang, Yujiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
by: Wu, Yinghao, et al.
Published: (2026)
by: Wu, Yinghao, et al.
Published: (2026)
Exploring Contextual Attribute Density in Referring Expression Counting
by: Wang, Zhicheng, et al.
Published: (2025)
by: Wang, Zhicheng, et al.
Published: (2025)
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
by: Luo, Zhuoyan, et al.
Published: (2024)
by: Luo, Zhuoyan, et al.
Published: (2024)
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
by: Chen, Qi, et al.
Published: (2025)
by: Chen, Qi, et al.
Published: (2025)
ML-SceGen: A Multi-level Scenario Generation Framework
by: Xiao, Yicheng, et al.
Published: (2025)
by: Xiao, Yicheng, et al.
Published: (2025)
Scalable Image Tokenization with Index Backpropagation Quantization
by: Shi, Fengyuan, et al.
Published: (2024)
by: Shi, Fengyuan, et al.
Published: (2024)
3D-GRES: Generalized 3D Referring Expression Segmentation
by: Wu, Changli, et al.
Published: (2024)
by: Wu, Changli, et al.
Published: (2024)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework
by: Zhang, Suoxiang, et al.
Published: (2025)
by: Zhang, Suoxiang, et al.
Published: (2025)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
by: Wang, Jingchao, et al.
Published: (2025)
by: Wang, Jingchao, et al.
Published: (2025)
Decoupling What to Count and Where to See for Referring Expression Counting
by: Zou, Yuda, et al.
Published: (2025)
by: Zou, Yuda, et al.
Published: (2025)
Decoding AI's Nudge: A Unified Framework to Predict Human Behavior in AI-assisted Decision Making
by: Li, Zhuoyan, et al.
Published: (2024)
by: Li, Zhuoyan, et al.
Published: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
by: Deng, Jingyuan, et al.
Published: (2025)
by: Deng, Jingyuan, et al.
Published: (2025)
Generalized Referring Expression Segmentation on Aerial Photos
by: Marnoto, Luís, et al.
Published: (2025)
by: Marnoto, Luís, et al.
Published: (2025)
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
by: Ding, Henghui, et al.
Published: (2026)
by: Ding, Henghui, et al.
Published: (2026)
Improving Contrastive Learning for Referring Expression Counting
by: Triaridis, Kostas, et al.
Published: (2025)
by: Triaridis, Kostas, et al.
Published: (2025)
Hierarchical Collaborative Fusion for 3D Instance-aware Referring Expression Segmentation
by: Zhou, Keshen, et al.
Published: (2026)
by: Zhou, Keshen, et al.
Published: (2026)
WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Latent Expression Generation for Referring Image Segmentation and Grounding
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
OpenCarbonEval: A Unified Carbon Emission Estimation Framework in Large-Scale AI Models
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
Pipelined Decoder for Efficient Context-Aware Text Generation
by: Huang, Zixian, et al.
Published: (2025)
by: Huang, Zixian, et al.
Published: (2025)
RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
by: Wu, Changli, et al.
Published: (2024)
by: Wu, Changli, et al.
Published: (2024)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
Alternative Approaches for Counting Weakly Increasing Matrices
by: Yang, Leo Yicheng
Published: (2025)
by: Yang, Leo Yicheng
Published: (2025)
ShareCMP: Polarization-Aware RGB-P Semantic Segmentation
by: Liu, Zhuoyan, et al.
Published: (2023)
by: Liu, Zhuoyan, et al.
Published: (2023)
Point, Segment and Count: A Generalized Framework for Object Counting
by: Huang, Zhizhong, et al.
Published: (2023)
by: Huang, Zhizhong, et al.
Published: (2023)
How Numerical Latitude Origin Expressions Increase Consumers' Purchase Intention: The Role of Perceived Naturalness
by: Minxue Huang, et al.
Published: (2026)
by: Minxue Huang, et al.
Published: (2026)
Bring Adaptive Binding Prototypes to Generalized Referring Expression Segmentation
by: Li, Weize, et al.
Published: (2024)
by: Li, Weize, et al.
Published: (2024)
Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation
by: Zhang, Weiming, et al.
Published: (2026)
by: Zhang, Weiming, et al.
Published: (2026)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
Progressive Knowledge Graph Completion
by: Li, Jiayi, et al.
Published: (2024)
by: Li, Jiayi, et al.
Published: (2024)
Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
by: Wang, Yaxian, et al.
Published: (2025)
by: Wang, Yaxian, et al.
Published: (2025)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
by: Nie, Sihang, et al.
Published: (2025)
by: Nie, Sihang, et al.
Published: (2025)
MACD: Model-Aware Contrastive Decoding via Counterfactual Data
by: Xiao, Qixin
Published: (2026)
by: Xiao, Qixin
Published: (2026)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
by: Dong, Qihua, et al.
Published: (2025)
by: Dong, Qihua, et al.
Published: (2025)
Hierarchical Skip Decoding for Efficient Autoregressive Text Generation
by: Zhu, Yunqi, et al.
Published: (2024)
by: Zhu, Yunqi, et al.
Published: (2024)
LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
by: Shi, Xianglong, et al.
Published: (2025)
by: Shi, Xianglong, et al.
Published: (2025)
HierarchicalForecast: A Reference Framework for Hierarchical Forecasting in Python
by: Olivares, Kin G., et al.
Published: (2022)
by: Olivares, Kin G., et al.
Published: (2022)
Similar Items
-
1st Place Solution for 5th LSVOS Challenge: Referring Video Object Segmentation
by: Luo, Zhuoyan, et al.
Published: (2024) -
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
by: Wu, Yinghao, et al.
Published: (2026) -
Exploring Contextual Attribute Density in Referring Expression Counting
by: Wang, Zhicheng, et al.
Published: (2025) -
Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
by: Luo, Zhuoyan, et al.
Published: (2024) -
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
by: Chen, Qi, et al.
Published: (2025)