CAMAv2: A Vision-Centric Approach for Static Map Element Annotation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Shiyuan, Zhang, Jiaxin, Mei, Ruohong, Cai, Yingfeng, Yin, Haoran, Chen, Tao, Sui, Wei, Yang, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Vision-Centric Approach for Static Map Element Annotation
by: Zhang, Jiaxin, et al.
Published: (2023)
by: Zhang, Jiaxin, et al.
Published: (2023)
VRSO: Visual-Centric Reconstruction for Static Object Annotation
by: Yu, Chenyao, et al.
Published: (2024)
by: Yu, Chenyao, et al.
Published: (2024)
RoMe: Towards Large Scale Road Surface Reconstruction via Mesh Representation
by: Mei, Ruohong, et al.
Published: (2023)
by: Mei, Ruohong, et al.
Published: (2023)
Unleashing Semantic and Geometric Priors for 3D Scene Completion
by: Chen, Shiyuan, et al.
Published: (2025)
by: Chen, Shiyuan, et al.
Published: (2025)
SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
by: Chen, Zhengqing, et al.
Published: (2025)
by: Chen, Zhengqing, et al.
Published: (2025)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
by: Jiao, Yang, et al.
Published: (2024)
by: Jiao, Yang, et al.
Published: (2024)
Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
by: Fang, Hongwei, et al.
Published: (2026)
by: Fang, Hongwei, et al.
Published: (2026)
Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
by: He, Weizhen, et al.
Published: (2025)
by: He, Weizhen, et al.
Published: (2025)
Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving
by: Yang, Yu, et al.
Published: (2024)
by: Yang, Yu, et al.
Published: (2024)
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models
by: Mei, Jianbiao, et al.
Published: (2025)
by: Mei, Jianbiao, et al.
Published: (2025)
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
by: Xie, Tao, et al.
Published: (2026)
by: Xie, Tao, et al.
Published: (2026)
Fine-Grained Zero-Shot Learning with Attribute-Centric Representations
by: Chen, Zhi, et al.
Published: (2025)
by: Chen, Zhi, et al.
Published: (2025)
MapExpert: Online HD Map Construction with Simple and Efficient Sparse Map Element Expert
by: Zhang, Dapeng, et al.
Published: (2024)
by: Zhang, Dapeng, et al.
Published: (2024)
Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Adaptive Anomaly Recovery for Telemanipulation: A Diffusion Model Approach to Vision-Based Tracking
by: Wang, Haoyang, et al.
Published: (2025)
by: Wang, Haoyang, et al.
Published: (2025)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
GSFusion: Online RGB-D Mapping Where Gaussian Splatting Meets TSDF Fusion
by: Wei, Jiaxin, et al.
Published: (2024)
by: Wei, Jiaxin, et al.
Published: (2024)
Fine-tuning Pre-trained Vision-Language Models in a Human-Annotation-Free Manner
by: Wang, Qian-Wei, et al.
Published: (2026)
by: Wang, Qian-Wei, et al.
Published: (2026)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
A Benchmark for Vision-Centric HD Mapping by V2I Systems
by: Fan, Miao, et al.
Published: (2025)
by: Fan, Miao, et al.
Published: (2025)
Static Key Attention in Vision
by: Hu, Zizhao, et al.
Published: (2024)
by: Hu, Zizhao, et al.
Published: (2024)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
FuncGrasp: Learning Object-Centric Neural Grasp Functions from Single Annotated Example Object
by: Chen, Hanzhi, et al.
Published: (2024)
by: Chen, Hanzhi, et al.
Published: (2024)
InPK: Infusing Prior Knowledge into Prompt for Vision-Language Models
by: Zhou, Shuchang, et al.
Published: (2025)
by: Zhou, Shuchang, et al.
Published: (2025)
Pre-Trained Vision-Language Models as Partial Annotators
by: Wang, Qian-Wei, et al.
Published: (2024)
by: Wang, Qian-Wei, et al.
Published: (2024)
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
by: Chen, Yin, et al.
Published: (2024)
by: Chen, Yin, et al.
Published: (2024)
Learning Sequence Descriptor based on Spatio-Temporal Attention for Visual Place Recognition
by: Zhao, Junqiao, et al.
Published: (2023)
by: Zhao, Junqiao, et al.
Published: (2023)
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
by: Zhu, Lei, et al.
Published: (2026)
by: Zhu, Lei, et al.
Published: (2026)
VIVECaption: A Split Approach to Caption Quality Improvement
by: Ananth, Varun, et al.
Published: (2026)
by: Ananth, Varun, et al.
Published: (2026)
Collaborative Multi-Mode Pruning for Vision-Language Models
by: Wu, Zimeng, et al.
Published: (2026)
by: Wu, Zimeng, et al.
Published: (2026)
Data Augmentation in Human-Centric Vision
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
RoDUS: Robust Decomposition of Static and Dynamic Elements in Urban Scenes
by: Nguyen, Thang-Anh-Quan, et al.
Published: (2024)
by: Nguyen, Thang-Anh-Quan, et al.
Published: (2024)
A Cross-Domain Few-Shot Learning Method Based on Domain Knowledge Mapping
by: Chen, Jiajun, et al.
Published: (2025)
by: Chen, Jiajun, et al.
Published: (2025)
Beyond Visual Cues: Synchronously Exploring Target-Centric Semantics for Vision-Language Tracking
by: Ge, Jiawei, et al.
Published: (2023)
by: Ge, Jiawei, et al.
Published: (2023)
Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations
by: Cui, Yibo, et al.
Published: (2025)
by: Cui, Yibo, et al.
Published: (2025)
Head Similarity: Modeling Structured Whole-Head Appearance Beyond Face Recognition
by: Wang, Yingfeng, et al.
Published: (2026)
by: Wang, Yingfeng, et al.
Published: (2026)
An Efficient Remote Sensing Super Resolution Method Exploring Diffusion Priors and Multi-Modal Constraints for Crop Type Mapping
by: Yang, Songxi, et al.
Published: (2025)
by: Yang, Songxi, et al.
Published: (2025)
Fundus2Video: Cross-Modal Angiography Video Generation from Static Fundus Photography with Clinical Knowledge Guidance
by: Zhang, Weiyi, et al.
Published: (2024)
by: Zhang, Weiyi, et al.
Published: (2024)
Similar Items
-
A Vision-Centric Approach for Static Map Element Annotation
by: Zhang, Jiaxin, et al.
Published: (2023) -
VRSO: Visual-Centric Reconstruction for Static Object Annotation
by: Yu, Chenyao, et al.
Published: (2024) -
RoMe: Towards Large Scale Road Surface Reconstruction via Mesh Representation
by: Mei, Ruohong, et al.
Published: (2023) -
Unleashing Semantic and Geometric Priors for 3D Scene Completion
by: Chen, Shiyuan, et al.
Published: (2025) -
SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
by: Chen, Zhengqing, et al.
Published: (2025)