LVLM-Aided Alignment of Task-Specific Vision Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Koebler, Alexander, Kuhn, Lukas, Thon, Ingo, Buettner, Florian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Grasping Partially Occluded Objects Using Autoencoder-Based Point Cloud Inpainting
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
von: Kuhn, Lukas, et al.
Veröffentlicht: (2026)
von: Kuhn, Lukas, et al.
Veröffentlicht: (2026)
Incremental Uncertainty-aware Performance Monitoring with Active Labeling Intervention
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
von: Qharabagh, Muhammad Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Muhammad Fetrat, et al.
Veröffentlicht: (2024)
MoRE-LLM: Mixture of Rule Experts Guided by a Large Language Model
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)
Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
von: Qian, Jiaye, et al.
Veröffentlicht: (2025)
von: Qian, Jiaye, et al.
Veröffentlicht: (2025)
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
von: Buettner, Kyle, et al.
Veröffentlicht: (2025)
von: Buettner, Kyle, et al.
Veröffentlicht: (2025)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
von: Park, Seongheon, et al.
Veröffentlicht: (2026)
von: Park, Seongheon, et al.
Veröffentlicht: (2026)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
von: Buettner, Kyle, et al.
Veröffentlicht: (2024)
von: Buettner, Kyle, et al.
Veröffentlicht: (2024)
Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM
von: Li, Shen, et al.
Veröffentlicht: (2025)
von: Li, Shen, et al.
Veröffentlicht: (2025)
OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
von: Miyamoto, Ryoto, et al.
Veröffentlicht: (2025)
von: Miyamoto, Ryoto, et al.
Veröffentlicht: (2025)
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2026)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2026)
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
von: Zhao, Yaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Yaqi, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
von: Wu, Aodi, et al.
Veröffentlicht: (2025)
von: Wu, Aodi, et al.
Veröffentlicht: (2025)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
von: Li, Jiaming, et al.
Veröffentlicht: (2025)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
von: Mao, Jiawei, et al.
Veröffentlicht: (2026)
von: Mao, Jiawei, et al.
Veröffentlicht: (2026)
Free$^2$Guide: Training-Free Text-to-Video Alignment using Image LVLM
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding
von: Cho, Beomsik, et al.
Veröffentlicht: (2025)
von: Cho, Beomsik, et al.
Veröffentlicht: (2025)
NuWa: Deriving Lightweight Task-Specific Vision Transformers for Edge Devices
von: Wei, Ziteng, et al.
Veröffentlicht: (2025)
von: Wei, Ziteng, et al.
Veröffentlicht: (2025)
Safety Alignment for Vision Language Models
von: Liu, Zhendong, et al.
Veröffentlicht: (2024)
von: Liu, Zhendong, et al.
Veröffentlicht: (2024)
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
von: Huang, Yiyang, et al.
Veröffentlicht: (2025)
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2025)
CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code Generation
von: Doris, Anna C., et al.
Veröffentlicht: (2025)
von: Doris, Anna C., et al.
Veröffentlicht: (2025)
OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis
von: Lin, Tianwei, et al.
Veröffentlicht: (2026)
von: Lin, Tianwei, et al.
Veröffentlicht: (2026)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
von: Feng, Mingqian, et al.
Veröffentlicht: (2024)
von: Feng, Mingqian, et al.
Veröffentlicht: (2024)
Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
von: Kwon, JuneHyoung, et al.
Veröffentlicht: (2026)
Task-Specific Adaptation of Segmentation Foundation Model via Prompt Learning
von: Kim, Hyung-Il, et al.
Veröffentlicht: (2024)
von: Kim, Hyung-Il, et al.
Veröffentlicht: (2024)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
von: Saini, Harshvardhan, et al.
Veröffentlicht: (2026)
Learning to Look: Cognitive Attention Alignment with Vision-Language Models
von: Yang, Ryan L., et al.
Veröffentlicht: (2025)
von: Yang, Ryan L., et al.
Veröffentlicht: (2025)
Token-Level Inference-Time Alignment for Vision-Language Models
von: Chen, Kejia, et al.
Veröffentlicht: (2025)
von: Chen, Kejia, et al.
Veröffentlicht: (2025)
Subspace Alignment for Vision-Language Model Test-time Adaptation
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
Tuning Vision-Language Models with Candidate Labels by Prompt Alignment
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
Provably Better Explanations with Optimized Aggregation of Feature Attributions
von: Decker, Thomas, et al.
Veröffentlicht: (2024)
von: Decker, Thomas, et al.
Veröffentlicht: (2024)
Enhancing Medical Large Vision-Language Models via Alignment Distillation
von: Chang, Aofei, et al.
Veröffentlicht: (2025)
von: Chang, Aofei, et al.
Veröffentlicht: (2025)
Make Your LVLM KV Cache More Lightweight
von: Chen, Xihao, et al.
Veröffentlicht: (2026)
von: Chen, Xihao, et al.
Veröffentlicht: (2026)
Enhance Vision-Language Alignment with Noise
von: Huang, Sida, et al.
Veröffentlicht: (2024)
von: Huang, Sida, et al.
Veröffentlicht: (2024)
General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
von: Lange, Bernard, et al.
Veröffentlicht: (2025)
von: Lange, Bernard, et al.
Veröffentlicht: (2025)
Zero-Training Task-Specific Model Synthesis for Few-Shot Medical Image Classification
von: Qin, Yao, et al.
Veröffentlicht: (2025)
von: Qin, Yao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Grasping Partially Occluded Objects Using Autoencoder-Based Point Cloud Inpainting
von: Koebler, Alexander, et al.
Veröffentlicht: (2025) -
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
von: Kuhn, Lukas, et al.
Veröffentlicht: (2026) -
Incremental Uncertainty-aware Performance Monitoring with Active Labeling Intervention
von: Koebler, Alexander, et al.
Veröffentlicht: (2025) -
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
von: Qharabagh, Muhammad Fetrat, et al.
Veröffentlicht: (2024) -
MoRE-LLM: Mixture of Rule Experts Guided by a Large Language Model
von: Koebler, Alexander, et al.
Veröffentlicht: (2025)