Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Qi, Chen, Jixuan, Zhang, Kaiyi, Yu, Xinquan, Chan, Antoni B., Huang, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mahalanobis Distance-based Multi-view Optimal Transport for Multi-view Crowd Localization
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Active View Selection for Scene-level Multi-view Crowd Counting and Localization with Limited Labels
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
by: Zhang, Qi, et al.
Published: (2020)
by: Zhang, Qi, et al.
Published: (2020)
Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Semi-Supervised Multi-View Crowd Counting by Ranking Multi-View Fusion Models
by: Zhang, Qi, et al.
Published: (2025)
by: Zhang, Qi, et al.
Published: (2025)
Density-based Object Detection in Crowded Scenes
by: Zhao, Chenyang, et al.
Published: (2025)
by: Zhao, Chenyang, et al.
Published: (2025)
SynMVCrowd: A Large Synthetic Benchmark for Multi-view Crowd Counting and Localization
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View Selection
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
DynamicTrack: Advancing Gigapixel Tracking in Crowded Scenes
by: Zhao, Yunqi, et al.
Published: (2024)
by: Zhao, Yunqi, et al.
Published: (2024)
CountFormer: Multi-View Crowd Counting Transformer
by: Mo, Hong, et al.
Published: (2024)
by: Mo, Hong, et al.
Published: (2024)
WSCF-MVCC: Weakly-supervised Calibration-free Multi-view Crowd Counting
by: Li, Bin, et al.
Published: (2025)
by: Li, Bin, et al.
Published: (2025)
Dense Point-to-Mask Optimization with Reinforced Point Selection for Crowd Instance Segmentation
by: Chen, Hongru, et al.
Published: (2026)
by: Chen, Hongru, et al.
Published: (2026)
Point-to-Region Loss for Semi-Supervised Point-Based Crowd Counting
by: Lin, Wei, et al.
Published: (2025)
by: Lin, Wei, et al.
Published: (2025)
EnvSocial-Diff: A Diffusion-Based Crowd Simulation Model with Environmental Conditioning and Individual-Group Interaction
by: Zhao, Bingxue, et al.
Published: (2026)
by: Zhao, Bingxue, et al.
Published: (2026)
CrossView-GS: Cross-view Gaussian Splatting For Large-scale Scene Reconstruction
by: Zhang, Chenhao, et al.
Published: (2025)
by: Zhang, Chenhao, et al.
Published: (2025)
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
by: Huang, Jiyang, et al.
Published: (2026)
by: Huang, Jiyang, et al.
Published: (2026)
Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
by: Wan, Jia, et al.
Published: (2024)
by: Wan, Jia, et al.
Published: (2024)
Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding
by: Yu, Xinquan, et al.
Published: (2025)
by: Yu, Xinquan, et al.
Published: (2025)
Learning Tracking Representations from Single Point Annotations
by: Wu, Qiangqiang, et al.
Published: (2024)
by: Wu, Qiangqiang, et al.
Published: (2024)
TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Multi-View Crowd Counting With Self-Supervised Learning
by: Mo, Hong, et al.
Published: (2025)
by: Mo, Hong, et al.
Published: (2025)
Embodied Crowd Counting
by: Long, Runling, et al.
Published: (2025)
by: Long, Runling, et al.
Published: (2025)
One-Shot Crowd Counting With Density Guidance For Scene Adaptation
by: Chen, Jiwei, et al.
Published: (2026)
by: Chen, Jiwei, et al.
Published: (2026)
PoI: A Filter to Extract Pixel of Interest from Novel Views for Scene Coordinate Regression
by: Li, Feifei, et al.
Published: (2025)
by: Li, Feifei, et al.
Published: (2025)
Optimized View and Geometry Distillation from Multi-view Diffuser
by: Zhang, Youjia, et al.
Published: (2023)
by: Zhang, Youjia, et al.
Published: (2023)
Crowd-SAM: SAM as a Smart Annotator for Object Detection in Crowded Scenes
by: Cai, Zhi, et al.
Published: (2024)
by: Cai, Zhi, et al.
Published: (2024)
IndoorCrowd: A Multi-Scene Dataset for Human Detection, Segmentation, and Tracking with an Automated Annotation Pipeline
by: Nae, Sebastian-Ion, et al.
Published: (2026)
by: Nae, Sebastian-Ion, et al.
Published: (2026)
DragScene: Interactive 3D Scene Editing with Single-view Drag Instructions
by: Gu, Chenghao, et al.
Published: (2024)
by: Gu, Chenghao, et al.
Published: (2024)
CrowdTrack: A Benchmark for Difficult Multiple Pedestrian Tracking in Real Scenarios
by: Fu, Teng, et al.
Published: (2025)
by: Fu, Teng, et al.
Published: (2025)
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
Analysis of Unstructured High-Density Crowded Scenes for Crowd Monitoring
by: Matov, Alexandre
Published: (2024)
by: Matov, Alexandre
Published: (2024)
ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition
by: Huang, Ronggang, et al.
Published: (2025)
by: Huang, Ronggang, et al.
Published: (2025)
3D Scene Change Modeling With Consistent Multi-View Aggregation
by: Zhou, Zirui, et al.
Published: (2025)
by: Zhou, Zirui, et al.
Published: (2025)
ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking
by: Ge, Jiawei, et al.
Published: (2026)
by: Ge, Jiawei, et al.
Published: (2026)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
by: Li, Xiangtai, et al.
Published: (2025)
by: Li, Xiangtai, et al.
Published: (2025)
ICG-MVSNet: Learning Intra-view and Cross-view Relationships for Guidance in Multi-View Stereo
by: Hu, Yuxi, et al.
Published: (2025)
by: Hu, Yuxi, et al.
Published: (2025)
Momentum-GS: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction
by: Fan, Jixuan, et al.
Published: (2024)
by: Fan, Jixuan, et al.
Published: (2024)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
MTMMC: A Large-Scale Real-World Multi-Modal Camera Tracking Benchmark
by: Woo, Sanghyun, et al.
Published: (2024)
by: Woo, Sanghyun, et al.
Published: (2024)
TrackRef3D: Multi-View Consistent Track-then-Label for Open-World Referring Segmentation in 3D Gaussian Splatting
by: Tan, Yuyang, et al.
Published: (2026)
by: Tan, Yuyang, et al.
Published: (2026)
Similar Items
-
Mahalanobis Distance-based Multi-view Optimal Transport for Multi-view Crowd Localization
by: Zhang, Qi, et al.
Published: (2024) -
Active View Selection for Scene-level Multi-view Crowd Counting and Localization with Limited Labels
by: Zhang, Qi, et al.
Published: (2025) -
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
by: Zhang, Qi, et al.
Published: (2020) -
Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting
by: Zhang, Qi, et al.
Published: (2024) -
Semi-Supervised Multi-View Crowd Counting by Ranking Multi-View Fusion Models
by: Zhang, Qi, et al.
Published: (2025)