The Promise of RL for Autoregressive Image Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmadi, Saba, Awal, Rabiul, Sikarwar, Ankur, Kazemnejad, Amirhossein, Luo, Ge Ya, Rodriguez, Juan A., Rajeswar, Sai, Reddy, Siva, Pal, Christopher, Krojer, Benno, Agrawal, Aishwarya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023)
by: Mañas, Oscar, et al.
Published: (2023)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023)
by: Zhang, Le, et al.
Published: (2023)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
Learning Action and Reasoning-Centric Image Editing from Videos and Simulations
by: Krojer, Benno, et al.
Published: (2024)
by: Krojer, Benno, et al.
Published: (2024)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
by: Awal, Rabiul, et al.
Published: (2025)
by: Awal, Rabiul, et al.
Published: (2025)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
by: Sikarwar, Ankur, et al.
Published: (2026)
by: Sikarwar, Ankur, et al.
Published: (2026)
VinePPO: Refining Credit Assignment in RL Training of LLMs
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
Benchmarking Vision Language Models for Cultural Understanding
by: Nayak, Shravan, et al.
Published: (2024)
by: Nayak, Shravan, et al.
Published: (2024)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
by: Yang, Qian, et al.
Published: (2026)
by: Yang, Qian, et al.
Published: (2026)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
by: Nayak, Shravan, et al.
Published: (2025)
by: Nayak, Shravan, et al.
Published: (2025)
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
by: Krojer, Benno, et al.
Published: (2026)
by: Krojer, Benno, et al.
Published: (2026)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
by: Aghajohari, Milad, et al.
Published: (2025)
by: Aghajohari, Milad, et al.
Published: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
by: Marjanović, Sara Vera, et al.
Published: (2025)
by: Marjanović, Sara Vera, et al.
Published: (2025)
GenRL: Multimodal-foundation world models for generalization in embodied agents
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
The Continuity of Oppression: State of Power in Iran from Reza Shah II to Ayatollah Khomeini
by: Himanshu Sikarwar
Published: (2026)
by: Himanshu Sikarwar
Published: (2026)
Beyond Mode-Seeking RL: Trajectory-Balance Post-Training for Diffusion Language Models
by: Ahmadi, Saba, et al.
Published: (2026)
by: Ahmadi, Saba, et al.
Published: (2026)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
by: Bechard, Patrice, et al.
Published: (2025)
by: Bechard, Patrice, et al.
Published: (2025)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
by: Rodriguez, Juan A., et al.
Published: (2025)
by: Rodriguez, Juan A., et al.
Published: (2025)
Paremia about the Earth in the Russian Language: Linguocultural Research
by: Sedigheh Kazemnejad DAHKAEI
Published: (2020)
by: Sedigheh Kazemnejad DAHKAEI
Published: (2020)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
by: Masry, Ahmed, et al.
Published: (2025)
by: Masry, Ahmed, et al.
Published: (2025)
Explanation of the Concept of Professional Engagement in Clinical Nurses in Iran: A Conventional Qualitative Content Analysis
by: Somayeh Rezaie, et al.
Published: (2025)
by: Somayeh Rezaie, et al.
Published: (2025)
Contributors to fatigue among nurses working in critical care units: A qualitative study
by: Reyhaneh Abbaszadeh, et al.
Published: (2024)
by: Reyhaneh Abbaszadeh, et al.
Published: (2024)
Differentiable Autoencoding Neural Operator for Interpretable and Integrable Latent Space Modeling
by: Viknesh, Siva, et al.
Published: (2025)
by: Viknesh, Siva, et al.
Published: (2025)
Discovering Failure Modes in Vision-Language Models using RL
by: Jain, Kanishk, et al.
Published: (2026)
by: Jain, Kanishk, et al.
Published: (2026)
Grounding Computer Use Agents on Human Demonstrations
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Power Evaluation of IOT Application Layer Protocols
by: Shahrokhi, Amirhossein, et al.
Published: (2024)
by: Shahrokhi, Amirhossein, et al.
Published: (2024)
Therefore I am. I Think
by: Esakkiraja, Esakkivel, et al.
Published: (2026)
by: Esakkiraja, Esakkivel, et al.
Published: (2026)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
by: Krojer, Benno, et al.
Published: (2025)
by: Krojer, Benno, et al.
Published: (2025)
Gray Swan Factory: Making Extreme Events from Ordinary Cyclones
by: Hakim, Gregory J., et al.
Published: (2026)
by: Hakim, Gregory J., et al.
Published: (2026)
DStruct2Design: Data and Benchmarks for Data Structure Driven Generative Floor Plan Design
by: Luo, Zhi Hao, et al.
Published: (2024)
by: Luo, Zhi Hao, et al.
Published: (2024)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning
by: Adaikkappan, Valliappan Chidambaram, et al.
Published: (2026)
by: Adaikkappan, Valliappan Chidambaram, et al.
Published: (2026)
Quantum Artificial Intelligence for Mission-Critical Systems: Foundations, Architectural Elements, and Future Directions
by: Sai, Siva, et al.
Published: (2025)
by: Sai, Siva, et al.
Published: (2025)
Ctrl-V: Higher Fidelity Video Generation with Bounding-Box Controlled Object Motion
by: Luo, Ge Ya, et al.
Published: (2024)
by: Luo, Ge Ya, et al.
Published: (2024)
Learning to See the Elephant in the Room: Self-Supervised Context Reasoning in Humans and AI
by: Liu, Xiao, et al.
Published: (2022)
by: Liu, Xiao, et al.
Published: (2022)
Similar Items
-
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024) -
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023) -
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
by: Zhang, Le, et al.
Published: (2023) -
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023) -
Learning Action and Reasoning-Centric Image Editing from Videos and Simulations
by: Krojer, Benno, et al.
Published: (2024)