Trustworthy Large Models in Vision: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Ziyan, Xu, Li, Liu, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vision Generalist Model: A Survey
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
Diffusion Models in Low-Level Vision: A Survey
von: He, Chunming, et al.
Veröffentlicht: (2024)
von: He, Chunming, et al.
Veröffentlicht: (2024)
Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models
von: Korkmaz, Cansu, et al.
Veröffentlicht: (2025)
von: Korkmaz, Cansu, et al.
Veröffentlicht: (2025)
A Survey on Vision Autoregressive Model
von: Jiang, Kai, et al.
Veröffentlicht: (2024)
von: Jiang, Kai, et al.
Veröffentlicht: (2024)
The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
von: Ai, Wei, et al.
Veröffentlicht: (2026)
von: Ai, Wei, et al.
Veröffentlicht: (2026)
Compound Expression Recognition via Large Vision-Language Models
von: Yu, Jun, et al.
Veröffentlicht: (2025)
von: Yu, Jun, et al.
Veröffentlicht: (2025)
Generative Physical AI in Vision: A Survey
von: Liu, Daochang, et al.
Veröffentlicht: (2025)
von: Liu, Daochang, et al.
Veröffentlicht: (2025)
Assessing Color Vision Test in Large Vision-language Models
von: Ye, Hongfei, et al.
Veröffentlicht: (2025)
von: Ye, Hongfei, et al.
Veröffentlicht: (2025)
On the Trustworthiness Landscape of State-of-the-art Generative Models: A Survey and Outlook
von: Fan, Mingyuan, et al.
Veröffentlicht: (2023)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2023)
Vision Language Models in Autonomous Driving: A Survey and Outlook
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
Efficient Diffusion Models for Vision: A Survey
von: Ulhaq, Anwaar, et al.
Veröffentlicht: (2022)
von: Ulhaq, Anwaar, et al.
Veröffentlicht: (2022)
FoPru: Focal Pruning for Efficient Large Vision-Language Models
von: Jiang, Lei, et al.
Veröffentlicht: (2024)
von: Jiang, Lei, et al.
Veröffentlicht: (2024)
3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
von: Liu, Chaohu, et al.
Veröffentlicht: (2025)
von: Liu, Chaohu, et al.
Veröffentlicht: (2025)
Efficient Multimodal Large Language Models: A Survey
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
von: Jin, Yizhang, et al.
Veröffentlicht: (2024)
A Survey on Mamba Architecture for Vision Applications
von: Ibrahim, Fady, et al.
Veröffentlicht: (2025)
von: Ibrahim, Fady, et al.
Veröffentlicht: (2025)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
Challenges and Trends in Egocentric Vision: A Survey
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
von: Yu, Qiucheng, et al.
Veröffentlicht: (2026)
von: Yu, Qiucheng, et al.
Veröffentlicht: (2026)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
von: Mohammadshirazi, Ahmad, et al.
Veröffentlicht: (2024)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models
von: Zhong, Jing, et al.
Veröffentlicht: (2025)
von: Zhong, Jing, et al.
Veröffentlicht: (2025)
UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models
von: Yin, Jun, et al.
Veröffentlicht: (2025)
von: Yin, Jun, et al.
Veröffentlicht: (2025)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
Probing Perceptual Constancy in Large Vision-Language Models
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
TSTMotion: Training-free Scene-aware Text-to-motion Generation
von: Guo, Ziyan, et al.
Veröffentlicht: (2025)
von: Guo, Ziyan, et al.
Veröffentlicht: (2025)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
Harnessing Large Vision and Language Models in Agriculture: A Review
von: Zhu, Hongyan, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyan, et al.
Veröffentlicht: (2024)
FILA: Fine-Grained Vision Language Models
von: Zhu, Shiding, et al.
Veröffentlicht: (2024)
von: Zhu, Shiding, et al.
Veröffentlicht: (2024)
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
von: Li, Shiyao, et al.
Veröffentlicht: (2024)
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
von: Zhang, Zhenguo, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenguo, et al.
Veröffentlicht: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
KKA: Improving Vision Anomaly Detection through Anomaly-related Knowledge from Large Language Models
von: Chen, Dong, et al.
Veröffentlicht: (2025)
von: Chen, Dong, et al.
Veröffentlicht: (2025)
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
Local Feature Matching Using Deep Learning: A Survey
von: Xu, Shibiao, et al.
Veröffentlicht: (2024)
von: Xu, Shibiao, et al.
Veröffentlicht: (2024)
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
Edge Deep Learning in Computer Vision and Medical Diagnostics: A Comprehensive Survey
von: Xu, Yiwen, et al.
Veröffentlicht: (2026)
von: Xu, Yiwen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Vision Generalist Model: A Survey
von: Wang, Ziyi, et al.
Veröffentlicht: (2025) -
Diffusion Models in Low-Level Vision: A Survey
von: He, Chunming, et al.
Veröffentlicht: (2024) -
Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models
von: Korkmaz, Cansu, et al.
Veröffentlicht: (2025) -
A Survey on Vision Autoregressive Model
von: Jiang, Kai, et al.
Veröffentlicht: (2024) -
The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
von: Ai, Wei, et al.
Veröffentlicht: (2026)