Embodied Crowd Counting
Fuente:
arXiv
Saved in:
| Main Authors: | Long, Runling, Wang, Yunlong, Wan, Jia, Deng, Xiang, Zhu, Xinting, Guan, Weili, Chan, Antoni B., Nie, Liqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
by: Wan, Jia, et al.
Published: (2024)
by: Wan, Jia, et al.
Published: (2024)
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
by: Huang, Jiyang, et al.
Published: (2026)
by: Huang, Jiyang, et al.
Published: (2026)
Density-based Object Detection in Crowded Scenes
by: Zhao, Chenyang, et al.
Published: (2025)
by: Zhao, Chenyang, et al.
Published: (2025)
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
by: Zhang, Qi, et al.
Published: (2020)
by: Zhang, Qi, et al.
Published: (2020)
Point-to-Region Loss for Semi-Supervised Point-Based Crowd Counting
by: Lin, Wei, et al.
Published: (2025)
by: Lin, Wei, et al.
Published: (2025)
Dense Point-to-Mask Optimization with Reinforced Point Selection for Crowd Instance Segmentation
by: Chen, Hongru, et al.
Published: (2026)
by: Chen, Hongru, et al.
Published: (2026)
Active View Selection for Scene-level Multi-view Crowd Counting and Localization with Limited Labels
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Semi-Supervised Multi-View Crowd Counting by Ranking Multi-View Fusion Models
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Single Domain Generalization for Crowd Counting
by: Peng, Zhuoxuan, et al.
Published: (2024)
by: Peng, Zhuoxuan, et al.
Published: (2024)
Video Individual Counting for Moving Drones
by: Fan, Yaowu, et al.
Published: (2025)
by: Fan, Yaowu, et al.
Published: (2025)
CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion
by: Li, Yanyu, et al.
Published: (2025)
by: Li, Yanyu, et al.
Published: (2025)
Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
by: Shen, Leyang, et al.
Published: (2024)
by: Shen, Leyang, et al.
Published: (2024)
A Fixed-Point Approach to Unified Prompt-Based Counting
by: Lin, Wei, et al.
Published: (2024)
by: Lin, Wei, et al.
Published: (2024)
Semi-Supervised Crowd Counting from Unlabeled Data
by: Duan, Haoran, et al.
Published: (2021)
by: Duan, Haoran, et al.
Published: (2021)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
by: Wang, Shaokun, et al.
Published: (2026)
by: Wang, Shaokun, et al.
Published: (2026)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
by: Lyu, Yibo, et al.
Published: (2025)
by: Lyu, Yibo, et al.
Published: (2025)
Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods
by: Fan, Yaowu, et al.
Published: (2026)
by: Fan, Yaowu, et al.
Published: (2026)
TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge
by: Zhang, Jinrong, et al.
Published: (2026)
by: Zhang, Jinrong, et al.
Published: (2026)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
by: He, Xusheng, et al.
Published: (2026)
by: He, Xusheng, et al.
Published: (2026)
R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
by: Zhang, Renshan, et al.
Published: (2024)
by: Zhang, Renshan, et al.
Published: (2024)
Mahalanobis Distance-based Multi-view Optimal Transport for Multi-view Crowd Localization
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Object-Shot Enhanced Grounding Network for Egocentric Video
by: Feng, Yisen, et al.
Published: (2025)
by: Feng, Yisen, et al.
Published: (2025)
Local Information Matters: A Rethink of Crowd Counting
by: Pan, Tianhang, et al.
Published: (2025)
by: Pan, Tianhang, et al.
Published: (2025)
OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
Beyond Quantity: Distribution-Aware Labeling for Visual Grounding
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
by: Wen, Haokun, et al.
Published: (2026)
by: Wen, Haokun, et al.
Published: (2026)
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
by: Wang, Yuchen, et al.
Published: (2025)
by: Wang, Yuchen, et al.
Published: (2025)
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
by: Zhang, Haoyu, et al.
Published: (2023)
by: Zhang, Haoyu, et al.
Published: (2023)
Crowded Video Individual Counting Informed by Social Grouping and Spatial-Temporal Displacement Priors
by: Lu, Hao, et al.
Published: (2026)
by: Lu, Hao, et al.
Published: (2026)
One-Shot Crowd Counting With Density Guidance For Scene Adaptation
by: Chen, Jiwei, et al.
Published: (2026)
by: Chen, Jiwei, et al.
Published: (2026)
FGENet: Fine-Grained Extraction Network for Congested Crowd Counting
by: Ma, Hao-Yuan, et al.
Published: (2024)
by: Ma, Hao-Yuan, et al.
Published: (2024)
Learning Discriminative Features for Crowd Counting
by: Chen, Yuehai, et al.
Published: (2023)
by: Chen, Yuehai, et al.
Published: (2023)
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
by: Zhang, Renshan, et al.
Published: (2025)
by: Zhang, Renshan, et al.
Published: (2025)
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge
by: Chen, Zhiwei, et al.
Published: (2026)
by: Chen, Zhiwei, et al.
Published: (2026)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
by: Fu, Zhiheng, et al.
Published: (2026)
by: Fu, Zhiheng, et al.
Published: (2026)
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
by: Chu, Qiaohui, et al.
Published: (2026)
by: Chu, Qiaohui, et al.
Published: (2026)
Density Estimation and Crowd Counting
by: Sunil, Balachandra Devarangadi, et al.
Published: (2025)
by: Sunil, Balachandra Devarangadi, et al.
Published: (2025)
Similar Items
-
Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
by: Wan, Jia, et al.
Published: (2024) -
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
by: Huang, Jiyang, et al.
Published: (2026) -
Density-based Object Detection in Crowded Scenes
by: Zhao, Chenyang, et al.
Published: (2025) -
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
by: Zhang, Qi, et al.
Published: (2020) -
Point-to-Region Loss for Semi-Supervised Point-Based Crowd Counting
by: Lin, Wei, et al.
Published: (2025)