RoScenes: A Large-scale Multi-view 3D Dataset for Roadside Perception
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhu, Xiaosu, Sheng, Hualian, Cai, Sijia, Deng, Bing, Yang, Shaopeng, Liang, Qiao, Chen, Ken, Gao, Lianli, Song, Jingkuan, Ye, Jieping |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CT3D++: Improving 3D Object Detection with Keypoint-induced Channel-wise Transformer
por: Sheng, Hualian, et al.
Publicado: (2024)
por: Sheng, Hualian, et al.
Publicado: (2024)
PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Models
por: Zhang, Jinhua, et al.
Publicado: (2024)
por: Zhang, Jinhua, et al.
Publicado: (2024)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
por: Li, Hao, et al.
Publicado: (2023)
por: Li, Hao, et al.
Publicado: (2023)
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
por: Wang, Xuanhan, et al.
Publicado: (2025)
por: Wang, Xuanhan, et al.
Publicado: (2025)
EchoShot: Multi-Shot Portrait Video Generation
por: Wang, Jiahao, et al.
Publicado: (2025)
por: Wang, Jiahao, et al.
Publicado: (2025)
AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
por: Wang, Jiahao, et al.
Publicado: (2026)
por: Wang, Jiahao, et al.
Publicado: (2026)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
por: Yang, Yuxiao, et al.
Publicado: (2025)
por: Yang, Yuxiao, et al.
Publicado: (2025)
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
por: Cai, Xiao, et al.
Publicado: (2024)
por: Cai, Xiao, et al.
Publicado: (2024)
ALF: Adaptive Label Finetuning for Scene Graph Generation
por: Chen, Qishen, et al.
Publicado: (2023)
por: Chen, Qishen, et al.
Publicado: (2023)
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
por: Su, Sitong, et al.
Publicado: (2023)
por: Su, Sitong, et al.
Publicado: (2023)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
por: Xing, Youguang, et al.
Publicado: (2025)
por: Xing, Youguang, et al.
Publicado: (2025)
Debiased Orthogonal Boundary-Driven Efficient Noise Mitigation
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
por: Wang, Xuanhan, et al.
Publicado: (2025)
por: Wang, Xuanhan, et al.
Publicado: (2025)
RoCo-Sim: Enhancing Roadside Collaborative Perception through Foreground Simulation
por: Du, Yuwen, et al.
Publicado: (2025)
por: Du, Yuwen, et al.
Publicado: (2025)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
por: Guo, Jiaqi, et al.
Publicado: (2024)
por: Guo, Jiaqi, et al.
Publicado: (2024)
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
por: Yin, Xiaoran, et al.
Publicado: (2025)
por: Yin, Xiaoran, et al.
Publicado: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
por: Liu, Ke, et al.
Publicado: (2025)
por: Liu, Ke, et al.
Publicado: (2025)
Any Target Can be Offense: Adversarial Example Generation via Generalized Latent Infection
por: Sun, Youheng, et al.
Publicado: (2024)
por: Sun, Youheng, et al.
Publicado: (2024)
RCooper: A Real-world Large-scale Dataset for Roadside Cooperative Perception
por: Hao, Ruiyang, et al.
Publicado: (2024)
por: Hao, Ruiyang, et al.
Publicado: (2024)
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
por: Cai, Xiao, et al.
Publicado: (2026)
por: Cai, Xiao, et al.
Publicado: (2026)
Informative Scene Graph Generation via Debiasing
por: Gao, Lianli, et al.
Publicado: (2023)
por: Gao, Lianli, et al.
Publicado: (2023)
Reliable Few-shot Learning under Dual Noises
por: Zhang, Ji, et al.
Publicado: (2025)
por: Zhang, Ji, et al.
Publicado: (2025)
DePT: Decoupled Prompt Tuning
por: Zhang, Ji, et al.
Publicado: (2023)
por: Zhang, Ji, et al.
Publicado: (2023)
AICL: Action In-Context Learning for Video Diffusion Model
por: Liu, Jianzhi, et al.
Publicado: (2024)
por: Liu, Jianzhi, et al.
Publicado: (2024)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
por: Lyu, Xinyu, et al.
Publicado: (2024)
por: Lyu, Xinyu, et al.
Publicado: (2024)
Beyond the Majority: Long-tail Imitation Learning for Robotic Manipulation
por: Zhu, Junhong, et al.
Publicado: (2026)
por: Zhu, Junhong, et al.
Publicado: (2026)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
por: Chen, Beitao, et al.
Publicado: (2025)
por: Chen, Beitao, et al.
Publicado: (2025)
CFReID: Continual Few-shot Person Re-Identification
por: Ni, Hao, et al.
Publicado: (2025)
por: Ni, Hao, et al.
Publicado: (2025)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
por: Chen, Beitao, et al.
Publicado: (2025)
por: Chen, Beitao, et al.
Publicado: (2025)
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
por: Luo, Xu, et al.
Publicado: (2026)
por: Luo, Xu, et al.
Publicado: (2026)
RoLID-11K: A Dashcam Dataset for Small-Object Roadside Litter Detection
por: Wu, Tao, et al.
Publicado: (2026)
por: Wu, Tao, et al.
Publicado: (2026)
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
por: Chen, Cheng, et al.
Publicado: (2024)
por: Chen, Cheng, et al.
Publicado: (2024)
FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
por: Yuan, Shengming, et al.
Publicado: (2025)
por: Yuan, Shengming, et al.
Publicado: (2025)
From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
por: Zhang, Ji, et al.
Publicado: (2023)
por: Zhang, Ji, et al.
Publicado: (2023)
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
por: Hong, Yu, et al.
Publicado: (2025)
por: Hong, Yu, et al.
Publicado: (2025)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
por: Wu, Shihan, et al.
Publicado: (2024)
por: Wu, Shihan, et al.
Publicado: (2024)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
por: Zhang, Ji, et al.
Publicado: (2025)
por: Zhang, Ji, et al.
Publicado: (2025)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
por: Guan, Runwei, et al.
Publicado: (2025)
por: Guan, Runwei, et al.
Publicado: (2025)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
por: Zhang, Haonan, et al.
Publicado: (2024)
por: Zhang, Haonan, et al.
Publicado: (2024)
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
por: Chen, Cheng, et al.
Publicado: (2026)
por: Chen, Cheng, et al.
Publicado: (2026)
Ejemplares similares
-
CT3D++: Improving 3D Object Detection with Keypoint-induced Channel-wise Transformer
por: Sheng, Hualian, et al.
Publicado: (2024) -
PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Models
por: Zhang, Jinhua, et al.
Publicado: (2024) -
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
por: Li, Hao, et al.
Publicado: (2023) -
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
por: Wang, Xuanhan, et al.
Publicado: (2025) -
EchoShot: Multi-Shot Portrait Video Generation
por: Wang, Jiahao, et al.
Publicado: (2025)