MTMMC: A Large-Scale Real-World Multi-Modal Camera Tracking Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Woo, Sanghyun, Park, Kwanyong, Shin, Inkyu, Kim, Myungchul, Kweon, In So |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
by: Kim, Myungchul, et al.
Published: (2026)
by: Kim, Myungchul, et al.
Published: (2026)
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
by: Kang, Minjun, et al.
Published: (2026)
by: Kang, Minjun, et al.
Published: (2026)
Drag4D: Align Your Motion with Text-Driven 3D Scene Generation
by: Kang, Minjun, et al.
Published: (2025)
by: Kang, Minjun, et al.
Published: (2025)
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
by: Shin, Inkyu, et al.
Published: (2024)
by: Shin, Inkyu, et al.
Published: (2024)
RoundaboutHD: High-Resolution Real-World Urban Environment Benchmark for Multi-Camera Vehicle Tracking
by: Lin, Yuqiang, et al.
Published: (2025)
by: Lin, Yuqiang, et al.
Published: (2025)
MrGS: Multi-modal Radiance Fields with 3D Gaussian Splatting for RGB-Thermal Novel View Synthesis
by: Kweon, Minseong, et al.
Published: (2025)
by: Kweon, Minseong, et al.
Published: (2025)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
by: Park, Kwanyong, et al.
Published: (2024)
by: Park, Kwanyong, et al.
Published: (2024)
All-day Depth Completion via Thermal-LiDAR Fusion
by: Kim, Janghyun, et al.
Published: (2025)
by: Kim, Janghyun, et al.
Published: (2025)
360 in the Wild: Dataset for Depth Prediction and View Synthesis
by: Park, Kibaek, et al.
Published: (2024)
by: Park, Kibaek, et al.
Published: (2024)
UDC-VIT: A Real-World Video Dataset for Under-Display Cameras
by: Ahn, Kyusu, et al.
Published: (2025)
by: Ahn, Kyusu, et al.
Published: (2025)
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
by: Lee, Yongjin, et al.
Published: (2024)
by: Lee, Yongjin, et al.
Published: (2024)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
Decomposition of Neural Discrete Representations for Large-Scale 3D Mapping
by: Park, Minseong, et al.
Published: (2024)
by: Park, Minseong, et al.
Published: (2024)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
by: An, Sojung, et al.
Published: (2025)
by: An, Sojung, et al.
Published: (2025)
MultiDepth: Multi-Sample Priors for Refining Monocular Metric Depth Estimations in Indoor Scenes
by: Byun, Sanghyun, et al.
Published: (2024)
by: Byun, Sanghyun, et al.
Published: (2024)
AnthroTAP: Learning Point Tracking with Real-World Motion
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
Enhancing Feature Tracking Reliability for Visual Navigation using Real-Time Safety Filter
by: Kim, Dabin, et al.
Published: (2025)
by: Kim, Dabin, et al.
Published: (2025)
ADNet: A Large-Scale and Extensible Multi-Domain Benchmark for Anomaly Detection Across 380 Real-World Categories
by: Ling, Hai, et al.
Published: (2025)
by: Ling, Hai, et al.
Published: (2025)
REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark
by: Wasserman, Navve, et al.
Published: (2025)
by: Wasserman, Navve, et al.
Published: (2025)
Complementary Random Masking for RGB-Thermal Semantic Segmentation
by: Shin, Ukcheol, et al.
Published: (2023)
by: Shin, Ukcheol, et al.
Published: (2023)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
by: Zhang, Chenshuang, et al.
Published: (2025)
by: Zhang, Chenshuang, et al.
Published: (2025)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
by: He, Ju, et al.
Published: (2023)
by: He, Ju, et al.
Published: (2023)
OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction
by: Lee, Junyoung, et al.
Published: (2026)
by: Lee, Junyoung, et al.
Published: (2026)
Multi-Modal Building Change Detection for Large-Scale Small Changes: Benchmark and Baseline
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
OceanSplat: Object-aware Gaussian Splatting with Trinocular View Consistency for Underwater Scene Reconstruction
by: Kweon, Minseong, et al.
Published: (2026)
by: Kweon, Minseong, et al.
Published: (2026)
Multi-Camera Worker Tracking in Logistics Warehouse Considering Wide-Angle Distortion
by: Mori, Yuki, et al.
Published: (2025)
by: Mori, Yuki, et al.
Published: (2025)
OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking
by: Liang, Haiji, et al.
Published: (2024)
by: Liang, Haiji, et al.
Published: (2024)
ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object
by: Zhang, Chenshuang, et al.
Published: (2024)
by: Zhang, Chenshuang, et al.
Published: (2024)
MCTR: Multi Camera Tracking Transformer
by: Niculescu-Mizil, Alexandru, et al.
Published: (2024)
by: Niculescu-Mizil, Alexandru, et al.
Published: (2024)
Robust Camera-to-Mocap Calibration and Verification for Large-Scale Multi-Camera Data Capture
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
Digital Scale: Open-Source On-Device BMI Estimation from Smartphone Camera Images Trained on a Large-Scale Real-World Dataset
by: Manichand, Frederik Rajiv, et al.
Published: (2025)
by: Manichand, Frederik Rajiv, et al.
Published: (2025)
Deeply Supervised Flow-Based Generative Models
by: Shin, Inkyu, et al.
Published: (2025)
by: Shin, Inkyu, et al.
Published: (2025)
EvLight++: Low-Light Video Enhancement with an Event Camera: A Large-Scale Real-World Dataset, Novel Method, and More
by: Chen, Kanghao, et al.
Published: (2024)
by: Chen, Kanghao, et al.
Published: (2024)
Scale Equalization for Multi-Level Feature Fusion
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
by: Li, Linfei, et al.
Published: (2026)
by: Li, Linfei, et al.
Published: (2026)
A Robust Deep Networks based Multi-Object MultiCamera Tracking System for City Scale Traffic
by: Zaman, Muhammad Imran, et al.
Published: (2025)
by: Zaman, Muhammad Imran, et al.
Published: (2025)
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
Asynchronous Multi-Object Tracking with an Event Camera
by: Apps, Angus, et al.
Published: (2025)
by: Apps, Angus, et al.
Published: (2025)
GS-EVT: Cross-Modal Event Camera Tracking based on Gaussian Splatting
by: Liu, Tao, et al.
Published: (2024)
by: Liu, Tao, et al.
Published: (2024)
Similar Items
-
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
by: Kim, Myungchul, et al.
Published: (2026) -
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
by: Kang, Minjun, et al.
Published: (2026) -
Drag4D: Align Your Motion with Text-Driven 3D Scene Generation
by: Kang, Minjun, et al.
Published: (2025) -
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
by: Shin, Inkyu, et al.
Published: (2024) -
RoundaboutHD: High-Resolution Real-World Urban Environment Benchmark for Multi-Camera Vehicle Tracking
by: Lin, Yuqiang, et al.
Published: (2025)