REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Fulin, Xiao, Wenyi, Chen, Bin, Din, Liang, Gan, Leilei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2026)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2026)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
von: Ba, Ying, et al.
Veröffentlicht: (2025)
von: Ba, Ying, et al.
Veröffentlicht: (2025)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
von: Yue, Xinli, et al.
Veröffentlicht: (2025)
von: Yue, Xinli, et al.
Veröffentlicht: (2025)
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
von: Li, Shufan, et al.
Veröffentlicht: (2024)
von: Li, Shufan, et al.
Veröffentlicht: (2024)
Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation
von: Lei, Sen, et al.
Veröffentlicht: (2024)
von: Lei, Sen, et al.
Veröffentlicht: (2024)
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
von: Liang, Tianyi, et al.
Veröffentlicht: (2024)
von: Liang, Tianyi, et al.
Veröffentlicht: (2024)
Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
von: He, Qingdong, et al.
Veröffentlicht: (2025)
von: He, Qingdong, et al.
Veröffentlicht: (2025)
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
von: Li, Yaqi, et al.
Veröffentlicht: (2025)
von: Li, Yaqi, et al.
Veröffentlicht: (2025)
Describe-to-Score: Text-Guided Efficient Image Complexity Assessment
von: Liu, Shipeng, et al.
Veröffentlicht: (2025)
von: Liu, Shipeng, et al.
Veröffentlicht: (2025)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization
von: Shi, Chuancheng, et al.
Veröffentlicht: (2025)
von: Shi, Chuancheng, et al.
Veröffentlicht: (2025)
VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
von: Wu, Tianhe, et al.
Veröffentlicht: (2025)
von: Wu, Tianhe, et al.
Veröffentlicht: (2025)
Pareto-Guided Optimal Transport for Multi-Reward Alignment
von: Ba, Ying, et al.
Veröffentlicht: (2026)
von: Ba, Ying, et al.
Veröffentlicht: (2026)
Identity-Preserving Text-to-Image Generation via Dual-Level Feature Decoupling and Expert-Guided Fusion
von: Chen, Kewen, et al.
Veröffentlicht: (2025)
von: Chen, Kewen, et al.
Veröffentlicht: (2025)
Efficient Pretraining Model based on Multi-Scale Local Visual Field Feature Reconstruction for PCB CT Image Element Segmentation
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Personalized Safety Alignment for Text-to-Image Diffusion Models
von: Lei, Yu, et al.
Veröffentlicht: (2025)
von: Lei, Yu, et al.
Veröffentlicht: (2025)
InstructEngine: Instruction-driven Text-to-Image Alignment
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
Cross-Modal Urban Sensing: Evaluating Sound-Vision Alignment Across Street-Level and Aerial Imagery
von: Chen, Pengyu, et al.
Veröffentlicht: (2025)
von: Chen, Pengyu, et al.
Veröffentlicht: (2025)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
von: Zhao, Shijie, et al.
Veröffentlicht: (2025)
von: Zhao, Shijie, et al.
Veröffentlicht: (2025)
Forgedit: Text Guided Image Editing via Learning and Forgetting
von: Zhang, Shiwen, et al.
Veröffentlicht: (2023)
von: Zhang, Shiwen, et al.
Veröffentlicht: (2023)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
von: Bai, Sule, et al.
Veröffentlicht: (2025)
von: Bai, Sule, et al.
Veröffentlicht: (2025)
Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation
von: Zhang, Wenchao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenchao, et al.
Veröffentlicht: (2025)
Multi-Text Guided Few-Shot Semantic Segmentation
von: Jiao, Qiang, et al.
Veröffentlicht: (2025)
von: Jiao, Qiang, et al.
Veröffentlicht: (2025)
ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning
von: Chen, Honghua, et al.
Veröffentlicht: (2026)
von: Chen, Honghua, et al.
Veröffentlicht: (2026)
Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models
von: Xia, Tao, et al.
Veröffentlicht: (2026)
von: Xia, Tao, et al.
Veröffentlicht: (2026)
Fine-Grained Behavior and Lane Constraints Guided Trajectory Prediction Method
von: Xiong, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiong, Wenyi, et al.
Veröffentlicht: (2025)
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
von: Jose, Cijo, et al.
Veröffentlicht: (2024)
von: Jose, Cijo, et al.
Veröffentlicht: (2024)
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
von: Wang, Yandi, et al.
Veröffentlicht: (2026)
von: Wang, Yandi, et al.
Veröffentlicht: (2026)
GeoR-Bench: Evaluating Geoscience Visual Reasoning
von: Zheng, Yushuo, et al.
Veröffentlicht: (2026)
von: Zheng, Yushuo, et al.
Veröffentlicht: (2026)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
von: Ma, Leilei, et al.
Veröffentlicht: (2024)
von: Ma, Leilei, et al.
Veröffentlicht: (2024)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
Reinforcing Multimodal Reasoning Against Visual Degradation
von: Liu, Rui, et al.
Veröffentlicht: (2026)
von: Liu, Rui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025) -
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2026) -
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
von: Ba, Ying, et al.
Veröffentlicht: (2025) -
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
von: Qian, Yusu, et al.
Veröffentlicht: (2025) -
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
von: Yue, Xinli, et al.
Veröffentlicht: (2025)