Improving Large Vision and Language Models by Learning from a Panel of Peers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hernandez, Jefferson, Shi, Jing, Jenni, Simon, Ordonez, Vicente, Kafle, Kushal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
von: Lu, Jianglin, et al.
Veröffentlicht: (2026)
von: Lu, Jianglin, et al.
Veröffentlicht: (2026)
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
von: Just, Hoang Anh, et al.
Veröffentlicht: (2025)
von: Just, Hoang Anh, et al.
Veröffentlicht: (2025)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
von: Yang, Ziyan, et al.
Veröffentlicht: (2022)
von: Yang, Ziyan, et al.
Veröffentlicht: (2022)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
von: Wu, Qiucheng, et al.
Veröffentlicht: (2026)
von: Wu, Qiucheng, et al.
Veröffentlicht: (2026)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
von: Hua, Hang, et al.
Veröffentlicht: (2024)
von: Hua, Hang, et al.
Veröffentlicht: (2024)
SCoRD: Subject-Conditional Relation Detection with Text-Augmented Data
von: Yang, Ziyan, et al.
Veröffentlicht: (2023)
von: Yang, Ziyan, et al.
Veröffentlicht: (2023)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
von: Akdemir, Kiymet, et al.
Veröffentlicht: (2025)
von: Akdemir, Kiymet, et al.
Veröffentlicht: (2025)
ViC-MAE: Self-Supervised Representation Learning from Images and Video with Contrastive Masked Autoencoders
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2023)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2023)
Generative Visual Instruction Tuning
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
Are Bias Mitigation Techniques for Deep Learning Effective?
von: Shrestha, Robik, et al.
Veröffentlicht: (2021)
von: Shrestha, Robik, et al.
Veröffentlicht: (2021)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
von: Slyman, Eric, et al.
Veröffentlicht: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
von: Qi, Daiqing, et al.
Veröffentlicht: (2025)
von: Qi, Daiqing, et al.
Veröffentlicht: (2025)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
von: Koo, Jaywon, et al.
Veröffentlicht: (2025)
von: Koo, Jaywon, et al.
Veröffentlicht: (2025)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
von: Slyman, Eric, et al.
Veröffentlicht: (2024)
von: Slyman, Eric, et al.
Veröffentlicht: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2024)
von: Ray, Sourjyadip, et al.
Veröffentlicht: (2024)
Agentic Discovery with Active Hypothesis Exploration for Visual Recognition
von: Koo, Jaywon, et al.
Veröffentlicht: (2026)
von: Koo, Jaywon, et al.
Veröffentlicht: (2026)
Improving Taxonomic Image-based Out-of-distribution Detection With DNA Barcodes
von: Impiö, Mikko, et al.
Veröffentlicht: (2024)
von: Impiö, Mikko, et al.
Veröffentlicht: (2024)
Modular Prompt Learning Improves Vision-Language Models
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
von: Huang, Zhenhan, et al.
Veröffentlicht: (2025)
Improving Large Vision-Language Models' Understanding for Flow Field Data
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
GViT: Representing Images as Gaussians for Visual Recognition
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
They're All Doctors: Synthesizing Diverse Counterfactuals to Mitigate Associative Bias
von: Magid, Salma Abdel, et al.
Veröffentlicht: (2024)
von: Magid, Salma Abdel, et al.
Veröffentlicht: (2024)
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation
von: Zhang, Qihui, et al.
Veröffentlicht: (2025)
von: Zhang, Qihui, et al.
Veröffentlicht: (2025)
Out-of-Distribution Detection Using Peer-Class Generated by Large Language Model
von: Huang, K, et al.
Veröffentlicht: (2024)
von: Huang, K, et al.
Veröffentlicht: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Improved Alignment of Modalities in Large Vision Language Models
von: Jangra, Kartik, et al.
Veröffentlicht: (2025)
von: Jangra, Kartik, et al.
Veröffentlicht: (2025)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models
von: Suo, Wei, et al.
Veröffentlicht: (2024)
von: Suo, Wei, et al.
Veröffentlicht: (2024)
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2023)
von: Jing, Liqiang, et al.
Veröffentlicht: (2023)
LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation
von: Li, Zhenshi, et al.
Veröffentlicht: (2024)
von: Li, Zhenshi, et al.
Veröffentlicht: (2024)
Does Peer Observation Help? Vision-Sharing Collaboration for Vision-Language Navigation
von: Jin, Qunchao, et al.
Veröffentlicht: (2026)
von: Jin, Qunchao, et al.
Veröffentlicht: (2026)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
Locality Alignment Improves Vision-Language Models
von: Covert, Ian, et al.
Veröffentlicht: (2024)
von: Covert, Ian, et al.
Veröffentlicht: (2024)
Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations
von: Singh, Jaisidh, et al.
Veröffentlicht: (2024)
von: Singh, Jaisidh, et al.
Veröffentlicht: (2024)
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
von: Yuan, Kun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
von: Lu, Jianglin, et al.
Veröffentlicht: (2026) -
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
von: Just, Hoang Anh, et al.
Veröffentlicht: (2025) -
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
von: Yang, Ziyan, et al.
Veröffentlicht: (2022) -
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
von: Wu, Qiucheng, et al.
Veröffentlicht: (2026) -
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
von: Hua, Hang, et al.
Veröffentlicht: (2024)