UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jun, Tan, Shuo, Sun, Zelong, Gu, Tiancheng, Zhao, Yongle, Feng, Ziyong, Yang, Kaicheng, Lu, Zhiwu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
von: Shen, Hengyu, et al.
Veröffentlicht: (2026)
von: Shen, Hengyu, et al.
Veröffentlicht: (2026)
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
RWKV-CLIP: A Robust Vision-Language Representation Learner
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
von: Chen, Zhichao, et al.
Veröffentlicht: (2026)
von: Chen, Zhichao, et al.
Veröffentlicht: (2026)
Multi-label Cluster Discrimination for Visual Representation Learning
von: An, Xiang, et al.
Veröffentlicht: (2024)
von: An, Xiang, et al.
Veröffentlicht: (2024)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
von: Yang, Kaicheng, et al.
Veröffentlicht: (2024)
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
von: Zheng, Yuanlei, et al.
Veröffentlicht: (2026)
von: Zheng, Yuanlei, et al.
Veröffentlicht: (2026)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
von: Xie, Yin, et al.
Veröffentlicht: (2024)
von: Xie, Yin, et al.
Veröffentlicht: (2024)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
von: Zhao, Xinping, et al.
Veröffentlicht: (2024)
von: Zhao, Xinping, et al.
Veröffentlicht: (2024)
1st Place Solution to the 1st SkatingVerse Challenge
von: Sun, Tao, et al.
Veröffentlicht: (2024)
von: Sun, Tao, et al.
Veröffentlicht: (2024)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
From Semantics to Pixels: Coarse-to-Fine Masked Autoencoders for Hierarchical Visual Understanding
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2026)
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2026)
Region-based Cluster Discrimination for Visual Representation Learning
von: Xie, Yin, et al.
Veröffentlicht: (2025)
von: Xie, Yin, et al.
Veröffentlicht: (2025)
Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2024)
von: Sun, Zelong, et al.
Veröffentlicht: (2024)
Say Cheese! Detail-Preserving Portrait Collection Generation via Natural Language Edits
von: Sun, Zelong, et al.
Veröffentlicht: (2026)
von: Sun, Zelong, et al.
Veröffentlicht: (2026)
Decoupled Global-Local Alignment for Improving Compositional Understanding
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
Mixture of Horizons in Action Chunking
von: Jing, Dong, et al.
Veröffentlicht: (2025)
von: Jing, Dong, et al.
Veröffentlicht: (2025)
Coarse-to-Fine: Progressive Image Compression for Semantically Hierarchical Classification
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
von: Xie, Yin, et al.
Veröffentlicht: (2025)
von: Xie, Yin, et al.
Veröffentlicht: (2025)
CTRL-RAG: Contrastive Likelihood Reward Based Reinforcement Learning for Context-Faithful RAG Models
von: Tan, Zhehao, et al.
Veröffentlicht: (2026)
von: Tan, Zhehao, et al.
Veröffentlicht: (2026)
From Coarse to Fine: Self-Adaptive Hierarchical Planning for LLM Agents
von: Tan, Haoran, et al.
Veröffentlicht: (2026)
von: Tan, Haoran, et al.
Veröffentlicht: (2026)
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
von: Hu, Yulan, et al.
Veröffentlicht: (2025)
von: Hu, Yulan, et al.
Veröffentlicht: (2025)
DocReward: A Document Reward Model for Structuring and Stylizing
von: Liu, Junpeng, et al.
Veröffentlicht: (2025)
von: Liu, Junpeng, et al.
Veröffentlicht: (2025)
Handling Multiple Hypotheses in Coarse-to-Fine Dense Image Matching
von: Vilain, Matthieu, et al.
Veröffentlicht: (2025)
von: Vilain, Matthieu, et al.
Veröffentlicht: (2025)
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
von: Zheng, Tianlu, et al.
Veröffentlicht: (2025)
von: Zheng, Tianlu, et al.
Veröffentlicht: (2025)
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
von: Shin, Joongmin, et al.
Veröffentlicht: (2026)
von: Shin, Joongmin, et al.
Veröffentlicht: (2026)
TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning
von: von Klinski, Maximilian, et al.
Veröffentlicht: (2026)
von: von Klinski, Maximilian, et al.
Veröffentlicht: (2026)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
ResWM: Residual-Action World Model for Visual RL
von: Zhang, Jseen, et al.
Veröffentlicht: (2026)
von: Zhang, Jseen, et al.
Veröffentlicht: (2026)
Coarse-to-fine crack cue for robust crack detection
von: Liu, Zelong, et al.
Veröffentlicht: (2025)
von: Liu, Zelong, et al.
Veröffentlicht: (2025)
Coarse-to-Fine Learning of Dynamic Causal Structures
von: Yang, Dezhi, et al.
Veröffentlicht: (2026)
von: Yang, Dezhi, et al.
Veröffentlicht: (2026)
SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization
von: Huang, Yongle, et al.
Veröffentlicht: (2025)
von: Huang, Yongle, et al.
Veröffentlicht: (2025)
Reinforcement Fine-Tuning for History-Aware Dense Retriever in RAG
von: Zhang, Yicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yicheng, et al.
Veröffentlicht: (2026)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
von: Cui, Siying, et al.
Veröffentlicht: (2024)
von: Cui, Siying, et al.
Veröffentlicht: (2024)
HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
von: Wu, Junhao, et al.
Veröffentlicht: (2025)
von: Wu, Junhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
von: Gu, Tiancheng, et al.
Veröffentlicht: (2025) -
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
von: Shen, Hengyu, et al.
Veröffentlicht: (2026) -
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024) -
RWKV-CLIP: A Robust Vision-Language Representation Learner
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024) -
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2025)