Saved in:
| Main Authors: | Li, Yiming, Li, Zhiheng, Chen, Nuo, Gong, Moonjun, Lyu, Zonglin, Wang, Zehong, Jiang, Peili, Feng, Chen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.09383 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memorize What Matters: Emergent Scene Decomposition from Multitraverse
by: Li, Yiming, et al.
Published: (2024)
by: Li, Yiming, et al.
Published: (2024)
SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving
by: Li, Yiming, et al.
Published: (2023)
by: Li, Yiming, et al.
Published: (2023)
Tell Me Where You Are: Multimodal LLMs Meet Place Recognition
by: Lyu, Zonglin, et al.
Published: (2024)
by: Lyu, Zonglin, et al.
Published: (2024)
LiDAR-based 4D Occupancy Completion and Forecasting
by: Liu, Xinhao, et al.
Published: (2023)
by: Liu, Xinhao, et al.
Published: (2023)
TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation
by: Lyu, Zonglin, et al.
Published: (2025)
by: Lyu, Zonglin, et al.
Published: (2025)
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
by: Xu, Peng, et al.
Published: (2025)
by: Xu, Peng, et al.
Published: (2025)
Frame Interpolation with Consecutive Brownian Bridge Diffusion
by: Lyu, Zonglin, et al.
Published: (2024)
by: Lyu, Zonglin, et al.
Published: (2024)
MARS: An Instance-aware, Modular and Realistic Simulator for Autonomous Driving
by: Wu, Zirui, et al.
Published: (2023)
by: Wu, Zirui, et al.
Published: (2023)
CPO: Condition Preference Optimization for Controllable Image Generation
by: Lyu, Zonglin, et al.
Published: (2025)
by: Lyu, Zonglin, et al.
Published: (2025)
Multimodal Model for Computational Pathology:Representation Learning and Image Compression
by: Wu, Peihang, et al.
Published: (2026)
by: Wu, Peihang, et al.
Published: (2026)
MARS: Multimodal Active Robotic Sensing for Articulated Characterization
by: Zeng, Hongliang, et al.
Published: (2024)
by: Zeng, Hongliang, et al.
Published: (2024)
Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
by: Liu, Xinxin, et al.
Published: (2026)
by: Liu, Xinxin, et al.
Published: (2026)
MARS: Memory-Enhanced Agents with Reflective Self-improvement
by: Liang, Xuechen, et al.
Published: (2025)
by: Liang, Xuechen, et al.
Published: (2025)
3D Object Visibility Prediction in Autonomous Driving
by: Luo, Chuanyu, et al.
Published: (2024)
by: Luo, Chuanyu, et al.
Published: (2024)
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
by: Chng, Yong Xien, et al.
Published: (2025)
by: Chng, Yong Xien, et al.
Published: (2025)
DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation
by: Guo, Jiazhe, et al.
Published: (2025)
by: Guo, Jiazhe, et al.
Published: (2025)
DAOS: A Multimodal In-cabin Behavior Monitoring with Driver Action-Object Synergy Dataset
by: Li, Yiming, et al.
Published: (2026)
by: Li, Yiming, et al.
Published: (2026)
Curriculum Dataset Distillation
by: Ma, Zhiheng, et al.
Published: (2024)
by: Ma, Zhiheng, et al.
Published: (2024)
D2E-An Autonomous Decision-making Dataset involving Driver States and Human Evaluation
by: Ke, Zehong, et al.
Published: (2024)
by: Ke, Zehong, et al.
Published: (2024)
CLM: Removing the GPU Memory Barrier for 3D Gaussian Splatting
by: Zhao, Hexu, et al.
Published: (2025)
by: Zhao, Hexu, et al.
Published: (2025)
Geo$^\textbf{2}$: Geometry-Guided Cross-view Geo-Localization and Image Synthesis
by: Zhang, Yancheng, et al.
Published: (2026)
by: Zhang, Yancheng, et al.
Published: (2026)
NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments
by: Pan, Taiyi, et al.
Published: (2024)
by: Pan, Taiyi, et al.
Published: (2024)
Event-based Tiny Object Detection: A Benchmark Dataset and Baseline
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
MARS: a Multimodal Alignment and Ranking System for Few-Shot Segmentation
by: Catalano, Nico, et al.
Published: (2025)
by: Catalano, Nico, et al.
Published: (2025)
Self-Supervised Place Recognition by Refining Temporal and Featural Pseudo Labels from Panoramic Data
by: Chen, Chao, et al.
Published: (2022)
by: Chen, Chao, et al.
Published: (2022)
MARS: Mesh AutoRegressive Model for 3D Shape Detailization
by: Gao, Jingnan, et al.
Published: (2025)
by: Gao, Jingnan, et al.
Published: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing
by: Ma, Yaohui, et al.
Published: (2024)
by: Ma, Yaohui, et al.
Published: (2024)
TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis
by: He, Wanggui, et al.
Published: (2024)
by: He, Wanggui, et al.
Published: (2024)
Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks
by: Wang, Lehan, et al.
Published: (2024)
by: Wang, Lehan, et al.
Published: (2024)
OVMR: Open-Vocabulary Recognition with Multi-Modal References
by: Ma, Zehong, et al.
Published: (2024)
by: Ma, Zehong, et al.
Published: (2024)
ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
by: Guo, Xianda, et al.
Published: (2025)
by: Guo, Xianda, et al.
Published: (2025)
Open-sourced Data Ecosystem in Autonomous Driving: the Present and Future
by: Li, Hongyang, et al.
Published: (2023)
by: Li, Hongyang, et al.
Published: (2023)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
by: Gong, Biao, et al.
Published: (2023)
by: Gong, Biao, et al.
Published: (2023)
GI-GS: Global Illumination Decomposition on Gaussian Splatting for Inverse Rendering
by: Chen, Hongze, et al.
Published: (2024)
by: Chen, Hongze, et al.
Published: (2024)
Multimodal Fusion via Self-Consistent Task-Gradient Fields
by: Xiong, Jiayu, et al.
Published: (2024)
by: Xiong, Jiayu, et al.
Published: (2024)
MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
by: Wu, Zonglin, et al.
Published: (2025)
by: Wu, Zonglin, et al.
Published: (2025)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Similar Items
-
Memorize What Matters: Emergent Scene Decomposition from Multitraverse
by: Li, Yiming, et al.
Published: (2024) -
SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving
by: Li, Yiming, et al.
Published: (2023) -
Tell Me Where You Are: Multimodal LLMs Meet Place Recognition
by: Lyu, Zonglin, et al.
Published: (2024) -
LiDAR-based 4D Occupancy Completion and Forecasting
by: Liu, Xinhao, et al.
Published: (2023) -
TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation
by: Lyu, Zonglin, et al.
Published: (2025)