Saved in:
| Main Authors: | Zhu, Lianghui, Fang, Yuxin, Liao, Bencheng, Wang, Shijie, Cheng, Tianheng, Huang, Zilong, Chen, Chen, Wei, Lai, Zeng, Yutao, Wang, Ya, Lin, Yi, Li, Yu, Wang, Xinggang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.15619 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
by: Liao, Bencheng, et al.
Published: (2025)
by: Liao, Bencheng, et al.
Published: (2025)
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
by: Zhu, Lianghui, et al.
Published: (2023)
by: Zhu, Lianghui, et al.
Published: (2023)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
by: Li, Yingyue, et al.
Published: (2025)
by: Li, Yingyue, et al.
Published: (2025)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
by: Tao, Hongyuan, et al.
Published: (2025)
by: Tao, Hongyuan, et al.
Published: (2025)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
by: Zeng, Lunbin, et al.
Published: (2025)
by: Zeng, Lunbin, et al.
Published: (2025)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
by: Zou, Jialv, et al.
Published: (2024)
by: Zou, Jialv, et al.
Published: (2024)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
by: Zhu, Lianghui, et al.
Published: (2025)
by: Zhu, Lianghui, et al.
Published: (2025)
Polar Parametrization for Vision-based Surround-View 3D Detection
by: Chen, Shaoyu, et al.
Published: (2022)
by: Chen, Shaoyu, et al.
Published: (2022)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
Occupancy as Set of Points
by: Shi, Yiang, et al.
Published: (2024)
by: Shi, Yiang, et al.
Published: (2024)
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
by: Li, Yongkang, et al.
Published: (2024)
by: Li, Yongkang, et al.
Published: (2024)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
by: Zhu, Lianghui, et al.
Published: (2023)
by: Zhu, Lianghui, et al.
Published: (2023)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
by: Li, Zongming, et al.
Published: (2025)
by: Li, Zongming, et al.
Published: (2025)
WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition
by: Zhu, Lianghui, et al.
Published: (2024)
by: Zhu, Lianghui, et al.
Published: (2024)
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
by: Jiang, Haoyi, et al.
Published: (2024)
by: Jiang, Haoyi, et al.
Published: (2024)
EVA-02: A Visual Representation for Neon Genesis
by: Fang, Yuxin, et al.
Published: (2023)
by: Fang, Yuxin, et al.
Published: (2023)
Dynamics of Spinning Binary at 2PM
by: Chen, Gang, et al.
Published: (2024)
by: Chen, Gang, et al.
Published: (2024)
ControlAR: Controllable Image Generation with Autoregressive Models
by: Li, Zongming, et al.
Published: (2024)
by: Li, Zongming, et al.
Published: (2024)
Rectified Sparse Attention
by: Sun, Yutao, et al.
Published: (2025)
by: Sun, Yutao, et al.
Published: (2025)
Adaptive Fusion of Single-View and Multi-View Depth for Autonomous Driving
by: Cheng, JunDa, et al.
Published: (2024)
by: Cheng, JunDa, et al.
Published: (2024)
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
Universal YOCO for Efficient Depth Scaling
by: Sun, Yutao, et al.
Published: (2026)
by: Sun, Yutao, et al.
Published: (2026)
Multi-Level Damage-Aware Graph Learning for Resilient UAV Swarm Networks
by: Lin, Huan, et al.
Published: (2025)
by: Lin, Huan, et al.
Published: (2025)
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
by: Chen, Sili, et al.
Published: (2025)
by: Chen, Sili, et al.
Published: (2025)
Systematic integral evaluation for spin-resummed binary dynamics
by: Chen, Gang, et al.
Published: (2024)
by: Chen, Gang, et al.
Published: (2024)
Similar Items
-
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
by: Liao, Bencheng, et al.
Published: (2024) -
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
by: Zhu, Lianghui, et al.
Published: (2024) -
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
by: Zhu, Lianghui, et al.
Published: (2024) -
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
by: Liao, Bencheng, et al.
Published: (2023) -
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
by: Liao, Bencheng, et al.
Published: (2025)