RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Yuanhuiyi, Zheng, Xu, Jiang, Lutao, Yan, Yibo, Zou, Xin, Zhou, Huiyu, Zhang, Linfeng, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025)
by: Zou, Xin, et al.
Published: (2025)
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
BrightDreamer: Generic 3D Gaussian Generative Framework for Fast Text-to-3D Synthesis
by: Jiang, Lutao, et al.
Published: (2024)
by: Jiang, Lutao, et al.
Published: (2024)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025)
by: Huo, Jiahao, et al.
Published: (2025)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Unlocking Speech Instruction Data Potential with Query Rewriting
by: Hei, Yonghua, et al.
Published: (2025)
by: Hei, Yonghua, et al.
Published: (2025)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024)
by: Li, Jungang, et al.
Published: (2024)
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next
by: Pan, Ye, et al.
Published: (2026)
by: Pan, Ye, et al.
Published: (2026)
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
by: Zhou, Jiazhou, et al.
Published: (2025)
by: Zhou, Jiazhou, et al.
Published: (2025)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
by: Zou, Xin, et al.
Published: (2024)
by: Zou, Xin, et al.
Published: (2024)
Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
by: Zheng, Kening, et al.
Published: (2024)
by: Zheng, Kening, et al.
Published: (2024)
OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
by: Zhong, Ding, et al.
Published: (2025)
by: Zhong, Ding, et al.
Published: (2025)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
by: Zhang, Jianghangfan, et al.
Published: (2025)
by: Zhang, Jianghangfan, et al.
Published: (2025)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Multi-view Hypergraph-based Contrastive Learning Model for Cold-Start Micro-video Recommendation
by: Lyu, Sisuo, et al.
Published: (2024)
by: Lyu, Sisuo, et al.
Published: (2024)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
CompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene Layout
by: Bai, Haotian, et al.
Published: (2023)
by: Bai, Haotian, et al.
Published: (2023)
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
by: Zhou, Guanyu, et al.
Published: (2024)
by: Zhou, Guanyu, et al.
Published: (2024)
DiMeR: Disentangled Mesh Reconstruction Model
by: Jiang, Lutao, et al.
Published: (2025)
by: Jiang, Lutao, et al.
Published: (2025)
PANORAMA: The Rise of Omnidirectional Vision in the Embodied AI Era
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
by: Cheng, Xin, et al.
Published: (2024)
by: Cheng, Xin, et al.
Published: (2024)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
by: Yu, Shi, et al.
Published: (2024)
by: Yu, Shi, et al.
Published: (2024)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
by: Jia, Sihang, et al.
Published: (2026)
by: Jia, Sihang, et al.
Published: (2026)
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
Centering the Value of Every Modality: Towards Efficient and Resilient Modality-agnostic Semantic Segmentation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding
by: Zhou, Jiazhou, et al.
Published: (2023)
by: Zhou, Jiazhou, et al.
Published: (2023)
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
Similar Items
-
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024) -
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
by: Zheng, Xu, et al.
Published: (2024) -
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2025) -
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025) -
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
by: Lyu, Yuanhuiyi, et al.
Published: (2026)