DPO Learning with LLMs-Judge Signal for Computer Use Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Man, Cobbley, David, Su, Xin, Rosenman, Shachar, Lal, Vasudev, Tseng, Shao-Yen, Howard, Phillip |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation
by: Rosenman, Shachar, et al.
Published: (2023)
by: Rosenman, Shachar, et al.
Published: (2023)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
by: Aflalo, Estelle, et al.
Published: (2024)
by: Aflalo, Estelle, et al.
Published: (2024)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
by: Xu, Xiao, et al.
Published: (2022)
by: Xu, Xiao, et al.
Published: (2022)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Probing the Representational Power of Sparse Autoencoders in Vision Models
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
by: Olson, Matthew Lyle, et al.
Published: (2025)
by: Olson, Matthew Lyle, et al.
Published: (2025)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
by: Madasu, Avinash, et al.
Published: (2025)
by: Madasu, Avinash, et al.
Published: (2025)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
by: Kundu, Souvik, et al.
Published: (2025)
by: Kundu, Souvik, et al.
Published: (2025)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
by: Hinck, Musashi, et al.
Published: (2024)
by: Hinck, Musashi, et al.
Published: (2024)
FastRM: An efficient and automatic explainability framework for multimodal generative models
by: Stan, Gabriela Ben-Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben-Melech, et al.
Published: (2024)
SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples
by: Howard, Phillip, et al.
Published: (2023)
by: Howard, Phillip, et al.
Published: (2023)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)
by: Madasu, Avinash, et al.
Published: (2023)
L-MAGIC: Language Model Assisted Generation of Images with Coherence
by: Cai, Zhipeng, et al.
Published: (2024)
by: Cai, Zhipeng, et al.
Published: (2024)
Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples
by: Howard, Phillip, et al.
Published: (2026)
by: Howard, Phillip, et al.
Published: (2026)
Region-Normalized DPO for Medical Image Segmentation under Noisy Judges
by: Kalisch, Hamza, et al.
Published: (2026)
by: Kalisch, Hamza, et al.
Published: (2026)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
by: Stan, Gabriela Ben Melech, et al.
Published: (2024)
Computer-Use Agents as Judges for Generative User Interface
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
by: Choubey, Prafulla Kumar, et al.
Published: (2024)
Cross-Cultural Value Awareness in Large Vision-Language Models
by: Howard, Phillip, et al.
Published: (2026)
by: Howard, Phillip, et al.
Published: (2026)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
by: Yu, Sungduk, et al.
Published: (2025)
by: Yu, Sungduk, et al.
Published: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
by: Yu, Sungduk, et al.
Published: (2024)
by: Yu, Sungduk, et al.
Published: (2024)
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
by: Shang, Shuyao, et al.
Published: (2025)
by: Shang, Shuyao, et al.
Published: (2025)
Why do LLaVA Vision-Language Models Reply to Images in English?
by: Hinck, Musashi, et al.
Published: (2024)
by: Hinck, Musashi, et al.
Published: (2024)
PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
by: Huang, Qihan, et al.
Published: (2024)
by: Huang, Qihan, et al.
Published: (2024)
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
by: Ahn, Daechul, et al.
Published: (2024)
by: Ahn, Daechul, et al.
Published: (2024)
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
by: Tang, Jiyang, et al.
Published: (2025)
by: Tang, Jiyang, et al.
Published: (2025)
SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning
by: Cheng, Xin, et al.
Published: (2026)
by: Cheng, Xin, et al.
Published: (2026)
Scaling Agents for Computer Use
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
Adaptive Vision-Language Model Routing for Computer Use Agents
by: Liu, Xunzhuo, et al.
Published: (2026)
by: Liu, Xunzhuo, et al.
Published: (2026)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
by: Xie, Yuxi, et al.
Published: (2024)
by: Xie, Yuxi, et al.
Published: (2024)
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
by: Yang, Zhihe, et al.
Published: (2025)
by: Yang, Zhihe, et al.
Published: (2025)
Learning from Online Videos at Inference Time for Computer-Use Agents
by: Liu, Yujian, et al.
Published: (2025)
by: Liu, Yujian, et al.
Published: (2025)
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
by: Liu, Zhaoyang, et al.
Published: (2025)
by: Liu, Zhaoyang, et al.
Published: (2025)
Watch and Learn: Learning to Use Computers from Online Videos
by: Song, Chan Hee, et al.
Published: (2025)
by: Song, Chan Hee, et al.
Published: (2025)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
Similar Items
-
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation
by: Rosenman, Shachar, et al.
Published: (2023) -
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
by: Aflalo, Estelle, et al.
Published: (2024) -
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
by: Ratzlaff, Neale, et al.
Published: (2024) -
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024) -
BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
by: Xu, Xiao, et al.
Published: (2022)