Saved in:
| Main Authors: | Xie, Peijin, Qian, Shun, Liu, Bingquan, Wang, Dexin, Sun, Lin, Zhang, Xiangzheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.07538 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
by: Xie, Peijin, et al.
Published: (2024)
by: Xie, Peijin, et al.
Published: (2024)
End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
by: Zheng, Qiaoyu, et al.
Published: (2025)
by: Zheng, Qiaoyu, et al.
Published: (2025)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
by: Li, Yinglu, et al.
Published: (2025)
by: Li, Yinglu, et al.
Published: (2025)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
by: Zhu, Wang, et al.
Published: (2023)
by: Zhu, Wang, et al.
Published: (2023)
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
by: Qian, Shun, et al.
Published: (2024)
by: Qian, Shun, et al.
Published: (2024)
Knowledge-based learning in Text-RAG and Image-RAG
by: Shim, Alexander, et al.
Published: (2026)
by: Shim, Alexander, et al.
Published: (2026)
DLAFormer: An End-to-End Transformer For Document Layout Analysis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
by: Tan, Xichen, et al.
Published: (2025)
by: Tan, Xichen, et al.
Published: (2025)
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
by: Zhu, Minjie, et al.
Published: (2025)
by: Zhu, Minjie, et al.
Published: (2025)
SDformer: Efficient End-to-End Transformer for Depth Completion
by: Qian, Jian, et al.
Published: (2024)
by: Qian, Jian, et al.
Published: (2024)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
by: Cao, Linhan, et al.
Published: (2026)
by: Cao, Linhan, et al.
Published: (2026)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
by: Li, Bingxin
Published: (2025)
by: Li, Bingxin
Published: (2025)
Fully Unified Motion Planning for End-to-End Autonomous Driving
by: Liu, Lin, et al.
Published: (2025)
by: Liu, Lin, et al.
Published: (2025)
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
by: Dong, Daxiang, et al.
Published: (2026)
by: Dong, Daxiang, et al.
Published: (2026)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
by: Zhai, Yukun, et al.
Published: (2023)
by: Zhai, Yukun, et al.
Published: (2023)
CREPE: Coordinate-Aware End-to-End Document Parser
by: Okamoto, Yamato, et al.
Published: (2024)
by: Okamoto, Yamato, et al.
Published: (2024)
Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
by: Sun, Xiatao, et al.
Published: (2025)
by: Sun, Xiatao, et al.
Published: (2025)
End-to-End Visual Autonomous Parking via Control-Aided Attention
by: Chen, Chao, et al.
Published: (2025)
by: Chen, Chao, et al.
Published: (2025)
SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
by: Sun, Wenchao, et al.
Published: (2024)
by: Sun, Wenchao, et al.
Published: (2024)
Polar R-CNN: End-to-End Lane Detection with Fewer Anchors
by: Wang, Shengqi, et al.
Published: (2024)
by: Wang, Shengqi, et al.
Published: (2024)
An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion Models
by: Qu, Wentao, et al.
Published: (2024)
by: Qu, Wentao, et al.
Published: (2024)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
CoProU-VO: Combining Projected Uncertainty for End-to-End Unsupervised Monocular Visual Odometry
by: Xie, Jingchao, et al.
Published: (2025)
by: Xie, Jingchao, et al.
Published: (2025)
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
by: Jian, Siyong, et al.
Published: (2026)
by: Jian, Siyong, et al.
Published: (2026)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
by: Wang, Qiuchen, et al.
Published: (2026)
by: Wang, Qiuchen, et al.
Published: (2026)
YOLOv10: Real-Time End-to-End Object Detection
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
by: Liao, Bencheng, et al.
Published: (2023)
by: Liao, Bencheng, et al.
Published: (2023)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
by: Xue, Junxiao, et al.
Published: (2024)
by: Xue, Junxiao, et al.
Published: (2024)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
by: Liu, Shuming, et al.
Published: (2023)
by: Liu, Shuming, et al.
Published: (2023)
Fose: Fusion of One-Step Diffusion and End-to-End Network for Pansharpening
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
by: Ahn, Young Jin, et al.
Published: (2024)
by: Ahn, Young Jin, et al.
Published: (2024)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025)
by: Zhang, Chen-Lin, et al.
Published: (2025)
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
by: Zhu, Feng, et al.
Published: (2026)
by: Zhu, Feng, et al.
Published: (2026)
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
by: Liao, Bencheng, et al.
Published: (2024)
by: Liao, Bencheng, et al.
Published: (2024)
Similar Items
-
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
by: Xie, Peijin, et al.
Published: (2024) -
End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
by: Zheng, Qiaoyu, et al.
Published: (2025) -
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
by: Li, Yinglu, et al.
Published: (2025) -
Efficient End-to-End Visual Document Understanding with Rationale Distillation
by: Zhu, Wang, et al.
Published: (2023) -
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
by: Qian, Shun, et al.
Published: (2024)