VisualDeltas: Learning Preferences from Visual Quality Perturbations
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Hailiang, Liu, Yihao, Guan, Shengyue, Li, Haoze, Li, Sujian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
por: Guan, Shengyue, et al.
Publicado: (2026)
por: Guan, Shengyue, et al.
Publicado: (2026)
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
por: Geng, Scott, et al.
Publicado: (2025)
por: Geng, Scott, et al.
Publicado: (2025)
ViPO: Visual Preference Optimization at Scale
por: Li, Ming, et al.
Publicado: (2026)
por: Li, Ming, et al.
Publicado: (2026)
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
por: Li, Yuting, et al.
Publicado: (2025)
por: Li, Yuting, et al.
Publicado: (2025)
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
por: Ding, Yihao, et al.
Publicado: (2026)
por: Ding, Yihao, et al.
Publicado: (2026)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
por: Ding, Yihao, et al.
Publicado: (2025)
por: Ding, Yihao, et al.
Publicado: (2025)
Visual Preference Optimization with Rubric Rewards
por: Yu, Ya-Qi, et al.
Publicado: (2026)
por: Yu, Ya-Qi, et al.
Publicado: (2026)
When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?
por: Mu, Tongzhou, et al.
Publicado: (2024)
por: Mu, Tongzhou, et al.
Publicado: (2024)
Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy
por: Gao, Shujian, et al.
Publicado: (2026)
por: Gao, Shujian, et al.
Publicado: (2026)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
por: Shen, Zhixuan, et al.
Publicado: (2024)
por: Shen, Zhixuan, et al.
Publicado: (2024)
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation
por: Li, Mengtian, et al.
Publicado: (2026)
por: Li, Mengtian, et al.
Publicado: (2026)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
por: Li, Yuanshuai, et al.
Publicado: (2025)
por: Li, Yuanshuai, et al.
Publicado: (2025)
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?
por: Lee, Chia-Hsuan, et al.
Publicado: (2026)
por: Lee, Chia-Hsuan, et al.
Publicado: (2026)
EERPD: Leveraging Emotion and Emotion Regulation for Improving Personality Detection
por: Li, Zheng, et al.
Publicado: (2024)
por: Li, Zheng, et al.
Publicado: (2024)
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
por: Luo, Wen, et al.
Publicado: (2024)
por: Luo, Wen, et al.
Publicado: (2024)
What External Knowledge is Preferred by LLMs? Characterizing and Exploring Chain of Evidence in Imperfect Context for Multi-Hop QA
por: Chang, Zhiyuan, et al.
Publicado: (2024)
por: Chang, Zhiyuan, et al.
Publicado: (2024)
Hidden in Plain Sight: Visual-to-Symbolic Analytical Solution Inference from Field Visualizations
por: Li, Pengze, et al.
Publicado: (2026)
por: Li, Pengze, et al.
Publicado: (2026)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
por: Guan, Shengyue, et al.
Publicado: (2025)
por: Guan, Shengyue, et al.
Publicado: (2025)
StyleGuard: Preventing Text-to-Image-Model-based Style Mimicry Attacks by Style Perturbations
por: Li, Yanjie, et al.
Publicado: (2025)
por: Li, Yanjie, et al.
Publicado: (2025)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
por: Song, Yifan, et al.
Publicado: (2024)
por: Song, Yifan, et al.
Publicado: (2024)
Visual Instruction Bottleneck Tuning
por: Oh, Changdae, et al.
Publicado: (2025)
por: Oh, Changdae, et al.
Publicado: (2025)
KUDA: Keypoints to Unify Dynamics Learning and Visual Prompting for Open-Vocabulary Robotic Manipulation
por: Liu, Zixian, et al.
Publicado: (2025)
por: Liu, Zixian, et al.
Publicado: (2025)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
por: Chen, Cong, et al.
Publicado: (2025)
por: Chen, Cong, et al.
Publicado: (2025)
CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions
por: Zhan, Yuliang, et al.
Publicado: (2026)
por: Zhan, Yuliang, et al.
Publicado: (2026)
Parallel Reinforcement Learning Simulation for Visual Quadrotor Navigation
por: Saunders, Jack, et al.
Publicado: (2022)
por: Saunders, Jack, et al.
Publicado: (2022)
AI-generated Image Quality Assessment in Visual Communication
por: Tian, Yu, et al.
Publicado: (2024)
por: Tian, Yu, et al.
Publicado: (2024)
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
por: Zhao, Hanyang, et al.
Publicado: (2025)
por: Zhao, Hanyang, et al.
Publicado: (2025)
VQA$^2$: Visual Question Answering for Video Quality Assessment
por: Jia, Ziheng, et al.
Publicado: (2024)
por: Jia, Ziheng, et al.
Publicado: (2024)
Investigating Anisotropy in Visual Grounding under Controlled Counterfactual Perturbations
por: Lombardo, Gabriele, et al.
Publicado: (2026)
por: Lombardo, Gabriele, et al.
Publicado: (2026)
BitAbuse: A Dataset of Visually Perturbed Texts for Defending Phishing Attacks
por: Lee, Hanyong, et al.
Publicado: (2025)
por: Lee, Hanyong, et al.
Publicado: (2025)
Adaptive Masking Enhances Visual Grounding
por: Jia, Sen, et al.
Publicado: (2024)
por: Jia, Sen, et al.
Publicado: (2024)
Reinforcement Learning for Chain of Thought Compression with One-Domain-to-All Generalization
por: Li, Hanyu, et al.
Publicado: (2025)
por: Li, Hanyu, et al.
Publicado: (2025)
Harmful Visual Content Manipulation Matters in Misinformation Detection Under Multimedia Scenarios
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
por: Zhang, Tianle, et al.
Publicado: (2024)
por: Zhang, Tianle, et al.
Publicado: (2024)
Retrieval-Augmented Fine-Tuning With Preference Optimization For Visual Program Generation
por: Kang, Deokhyung, et al.
Publicado: (2025)
por: Kang, Deokhyung, et al.
Publicado: (2025)
Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations
por: Fu, Minghao, et al.
Publicado: (2026)
por: Fu, Minghao, et al.
Publicado: (2026)
Cerberus: Efficient Inference with Adaptive Parallel Decoding and Sequential Knowledge Enhancement
por: Liu, Yuxuan, et al.
Publicado: (2024)
por: Liu, Yuxuan, et al.
Publicado: (2024)
Hit-RAG: Learning to Reason with Long Contexts via Preference Alignment
por: Liu, Junming, et al.
Publicado: (2026)
por: Liu, Junming, et al.
Publicado: (2026)
Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals
por: Li, Jia-Nan, et al.
Publicado: (2025)
por: Li, Jia-Nan, et al.
Publicado: (2025)
Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages
por: Ma, Guozheng, et al.
Publicado: (2023)
por: Ma, Guozheng, et al.
Publicado: (2023)
Ejemplares similares
-
SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management
por: Guan, Shengyue, et al.
Publicado: (2026) -
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
por: Geng, Scott, et al.
Publicado: (2025) -
ViPO: Visual Preference Optimization at Scale
por: Li, Ming, et al.
Publicado: (2026) -
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
por: Li, Yuting, et al.
Publicado: (2025) -
Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding
por: Ding, Yihao, et al.
Publicado: (2026)