UNICBench: UNIfied Counting Benchmark for MLLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rong, Chenggang, Han, Tao, Zhao, Zhiyuan, Fan, Yaowu, Wan, Jia, Guo, Song, Yuan, Yuan, Gao, Junyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods
von: Fan, Yaowu, et al.
Veröffentlicht: (2026)
von: Fan, Yaowu, et al.
Veröffentlicht: (2026)
Video Individual Counting for Moving Drones
von: Fan, Yaowu, et al.
Veröffentlicht: (2025)
von: Fan, Yaowu, et al.
Veröffentlicht: (2025)
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
von: Zhang, Da, et al.
Veröffentlicht: (2025)
von: Zhang, Da, et al.
Veröffentlicht: (2025)
NWPU-MOC: A Benchmark for Fine-grained Multi-category Object Counting in Aerial Images
von: Gao, Junyu, et al.
Veröffentlicht: (2024)
von: Gao, Junyu, et al.
Veröffentlicht: (2024)
UniVG: Towards UNIfied-modal Video Generation
von: Ruan, Ludan, et al.
Veröffentlicht: (2024)
von: Ruan, Ludan, et al.
Veröffentlicht: (2024)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
von: Zhang, Da, et al.
Veröffentlicht: (2026)
von: Zhang, Da, et al.
Veröffentlicht: (2026)
Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline
von: Li, Bingyu, et al.
Veröffentlicht: (2026)
von: Li, Bingyu, et al.
Veröffentlicht: (2026)
UNICON: UNIfied CONtinual Learning for Medical Foundational Models
von: Qazi, Mohammad Areeb, et al.
Veröffentlicht: (2025)
von: Qazi, Mohammad Areeb, et al.
Veröffentlicht: (2025)
Prototype-Based Low Altitude UAV Semantic Segmentation
von: Zhang, Da, et al.
Veröffentlicht: (2026)
von: Zhang, Da, et al.
Veröffentlicht: (2026)
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2026)
von: Li, Bingyu, et al.
Veröffentlicht: (2026)
Real-Time Crowd Counting for Embedded Systems with Lightweight Architecture
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Zhiyuan, et al.
Veröffentlicht: (2025)
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
Quantum-inspired Interpretable Deep Learning Architecture for Text Sentiment Analysis
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
Spotlight Text Detector: Spotlight on Candidate Regions Like a Camera
von: Han, Xu, et al.
Veröffentlicht: (2024)
von: Han, Xu, et al.
Veröffentlicht: (2024)
Real-Time Text Detection with Similar Mask in Traffic, Industrial, and Natural Scenes
von: Han, Xu, et al.
Veröffentlicht: (2024)
von: Han, Xu, et al.
Veröffentlicht: (2024)
Focus Entirety and Perceive Environment for Arbitrary-Shaped Text Detection
von: Han, Xu, et al.
Veröffentlicht: (2024)
von: Han, Xu, et al.
Veröffentlicht: (2024)
Guided Depth Map Super-Resolution via Multi-Scale Fusion U-shaped Mamba Network
von: Guo, Chenggang, et al.
Veröffentlicht: (2025)
von: Guo, Chenggang, et al.
Veröffentlicht: (2025)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
von: Xu, Jingwei, et al.
Veröffentlicht: (2024)
SamLP: A Customized Segment Anything Model for License Plate Detection
von: Ding, Haoxuan, et al.
Veröffentlicht: (2024)
von: Ding, Haoxuan, et al.
Veröffentlicht: (2024)
One-Shot Crowd Counting With Density Guidance For Scene Adaptation
von: Chen, Jiwei, et al.
Veröffentlicht: (2026)
von: Chen, Jiwei, et al.
Veröffentlicht: (2026)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
von: Zhu, Huilin, et al.
Veröffentlicht: (2025)
von: Zhu, Huilin, et al.
Veröffentlicht: (2025)
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
von: Guo, Meng-Hao, et al.
Veröffentlicht: (2025)
von: Guo, Meng-Hao, et al.
Veröffentlicht: (2025)
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
A New People-Object Interaction Dataset and NVS Benchmarks
von: Guo, Shuai, et al.
Veröffentlicht: (2024)
von: Guo, Shuai, et al.
Veröffentlicht: (2024)
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
von: Zhang, Yipeng, et al.
Veröffentlicht: (2024)
von: Zhang, Yipeng, et al.
Veröffentlicht: (2024)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
von: Li, Zongjian, et al.
Veröffentlicht: (2025)
von: Li, Zongjian, et al.
Veröffentlicht: (2025)
U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
Dynamic Proxy Domain Generalizes the Crowd Localization by Better Binary Segmentation
von: Gao, Junyu, et al.
Veröffentlicht: (2024)
von: Gao, Junyu, et al.
Veröffentlicht: (2024)
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations
von: Yuan, Jiangye, et al.
Veröffentlicht: (2026)
von: Yuan, Jiangye, et al.
Veröffentlicht: (2026)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
Efficient Motion-Aware Video MLLM
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
Exploring the Underwater World Segmentation without Extra Training
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
A Benchmark for Multi-Lingual Vision-Language Learning in Remote Sensing Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2025)
von: Zhou, Qing, et al.
Veröffentlicht: (2025)
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning
von: Zhao, Rui, et al.
Veröffentlicht: (2025)
von: Zhao, Rui, et al.
Veröffentlicht: (2025)
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation
von: Li, Zhiheng, et al.
Veröffentlicht: (2026)
von: Li, Zhiheng, et al.
Veröffentlicht: (2026)
Diffusion-based Data Augmentation for Object Counting Problems
von: Wang, Zhen, et al.
Veröffentlicht: (2024)
von: Wang, Zhen, et al.
Veröffentlicht: (2024)
Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation
von: Wu, Yebo, et al.
Veröffentlicht: (2026)
von: Wu, Yebo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods
von: Fan, Yaowu, et al.
Veröffentlicht: (2026) -
Video Individual Counting for Moving Drones
von: Fan, Yaowu, et al.
Veröffentlicht: (2025) -
UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
von: Zhang, Da, et al.
Veröffentlicht: (2025) -
NWPU-MOC: A Benchmark for Fine-grained Multi-category Object Counting in Aerial Images
von: Gao, Junyu, et al.
Veröffentlicht: (2024) -
UniVG: Towards UNIfied-modal Video Generation
von: Ruan, Ludan, et al.
Veröffentlicht: (2024)