Saved in:
| Main Authors: | Fan, Siqi, Xie, Yuguang, Cai, Bowen, Xie, Ailin, Liu, Gaochao, Qiao, Mu, Xing, Jie, Nie, Zaiqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2501.15415 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Cooperative Trajectory Representations for Motion Forecasting
by: Ruan, Hongzhi, et al.
Published: (2023)
by: Ruan, Hongzhi, et al.
Published: (2023)
QUEST: Query Stream for Practical Cooperative Perception
by: Fan, Siqi, et al.
Published: (2023)
by: Fan, Siqi, et al.
Published: (2023)
StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving
by: Hao, Ruiyang, et al.
Published: (2025)
by: Hao, Ruiyang, et al.
Published: (2025)
End-to-End Autonomous Driving through V2X Cooperation
by: Yu, Haibao, et al.
Published: (2024)
by: Yu, Haibao, et al.
Published: (2024)
BioMedGPT-Mol: Multi-task Learning for Molecular Understanding and Generation
by: Zuo, Chenyang, et al.
Published: (2025)
by: Zuo, Chenyang, et al.
Published: (2025)
Situational Scene Graph for Structured Human-centric Situation Understanding
by: Sugandhika, Chinthani, et al.
Published: (2024)
by: Sugandhika, Chinthani, et al.
Published: (2024)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
GVKF: Gaussian Voxel Kernel Functions for Highly Efficient Surface Reconstruction in Open Scenes
by: Song, Gaochao, et al.
Published: (2024)
by: Song, Gaochao, et al.
Published: (2024)
CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential Perception
by: Zhong, Jiaru, et al.
Published: (2025)
by: Zhong, Jiaru, et al.
Published: (2025)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control
by: Xie, Linxi, et al.
Published: (2026)
by: Xie, Linxi, et al.
Published: (2026)
RCooper: A Real-world Large-scale Dataset for Roadside Cooperative Perception
by: Hao, Ruiyang, et al.
Published: (2024)
by: Hao, Ruiyang, et al.
Published: (2024)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
by: Wang, Ziteng, et al.
Published: (2025)
by: Wang, Ziteng, et al.
Published: (2025)
DriveE2E: Closed-Loop Benchmark for End-to-End Autonomous Driving through Real-to-Simulation
by: Yu, Haibao, et al.
Published: (2025)
by: Yu, Haibao, et al.
Published: (2025)
Hypergraph-Enhanced Training-Free and Language-Free Few-Shot Anomaly Detection
by: Xie, Guohuan, et al.
Published: (2026)
by: Xie, Guohuan, et al.
Published: (2026)
Let the Abyss Stare Back Adaptive Falsification for Autonomous Scientific Discovery
by: Li, Peiran, et al.
Published: (2026)
by: Li, Peiran, et al.
Published: (2026)
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Vision Transformer with Sparse Scan Prior
by: Zhang, Yuguang, et al.
Published: (2024)
by: Zhang, Yuguang, et al.
Published: (2024)
Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection
by: Yang, Yuguang, et al.
Published: (2025)
by: Yang, Yuguang, et al.
Published: (2025)
A Unified Framework for Human-centric Point Cloud Video Understanding
by: Xu, Yiteng, et al.
Published: (2024)
by: Xu, Yiteng, et al.
Published: (2024)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
by: Xie, Wulin, et al.
Published: (2025)
by: Xie, Wulin, et al.
Published: (2025)
Physics-Guided Image Dehazing Diffusion
by: Zhou, Shijun, et al.
Published: (2025)
by: Zhou, Shijun, et al.
Published: (2025)
Flame quality monitoring of flare stack based on deep visual features
by: Mu, Xing
Published: (2024)
by: Mu, Xing
Published: (2024)
Controllable Human-centric Keyframe Interpolation with Generative Prior
by: Guo, Zujin, et al.
Published: (2025)
by: Guo, Zujin, et al.
Published: (2025)
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
3D Skew Gaussian Splatting with Any Camera Trajectory Visualization Engine
by: Zhao, Beizhen, et al.
Published: (2026)
by: Zhao, Beizhen, et al.
Published: (2026)
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
by: Qiao, Yanyuan, et al.
Published: (2025)
by: Qiao, Yanyuan, et al.
Published: (2025)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
Atom-Level Optical Chemical Structure Recognition with Limited Supervision
by: Oldenhof, Martijn, et al.
Published: (2024)
by: Oldenhof, Martijn, et al.
Published: (2024)
Vision-centric Token Compression in Large Language Model
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
An Integrated Neighborhood and Scale Information Network for Open-Pit Mine Change Detection in High-Resolution Remote Sensing Images
by: Xie, Zilin, et al.
Published: (2024)
by: Xie, Zilin, et al.
Published: (2024)
Graph-Guided Scene Reconstruction from Images with 3D Gaussian Splatting
by: Cheng, Chong, et al.
Published: (2025)
by: Cheng, Chong, et al.
Published: (2025)
ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments
by: Chen, Jiali, et al.
Published: (2025)
by: Chen, Jiali, et al.
Published: (2025)
Rethinking Iterative Stereo Matching from Diffusion Bridge Model Perspective
by: Shi, Yuguang
Published: (2024)
by: Shi, Yuguang
Published: (2024)
Learning Part Knowledge to Facilitate Category Understanding for Fine-Grained Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2025)
by: Wang, Enguang, et al.
Published: (2025)
Research Challenges and Progress in the End-to-End V2X Cooperative Autonomous Driving Competition
by: Hao, Ruiyang, et al.
Published: (2025)
by: Hao, Ruiyang, et al.
Published: (2025)
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
SeqGrowGraph: Learning Lane Topology as a Chain of Graph Expansions
by: Xie, Mengwei, et al.
Published: (2025)
by: Xie, Mengwei, et al.
Published: (2025)
Energy-Guided Decoding for Object Hallucination Mitigation
by: Liu, Xixi, et al.
Published: (2025)
by: Liu, Xixi, et al.
Published: (2025)
Similar Items
-
Learning Cooperative Trajectory Representations for Motion Forecasting
by: Ruan, Hongzhi, et al.
Published: (2023) -
QUEST: Query Stream for Practical Cooperative Perception
by: Fan, Siqi, et al.
Published: (2023) -
StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving
by: Hao, Ruiyang, et al.
Published: (2025) -
End-to-End Autonomous Driving through V2X Cooperation
by: Yu, Haibao, et al.
Published: (2024) -
BioMedGPT-Mol: Multi-task Learning for Molecular Understanding and Generation
by: Zuo, Chenyang, et al.
Published: (2025)