VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Ju, Jeongho, Kim, Daeyoung, Park, SunYoung, Kim, Youngjune |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VARCO-VISION-2.0 Technical Report
di: Cha, Young-rok, et al.
Pubblicazione: (2025)
di: Cha, Young-rok, et al.
Pubblicazione: (2025)
V-Agent: An Interactive Video Search System Using Vision-Language Models
di: Park, SunYoung, et al.
Pubblicazione: (2025)
di: Park, SunYoung, et al.
Pubblicazione: (2025)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
di: Kim, Yoonshik, et al.
Pubblicazione: (2025)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
di: Lee, Seongyun, et al.
Pubblicazione: (2024)
Evaluating Multimodal Generative AI with Korean Educational Standards
di: Park, Sanghee, et al.
Pubblicazione: (2025)
di: Park, Sanghee, et al.
Pubblicazione: (2025)
Reference-Based Post-OCR Processing with LLM for Precise Diacritic Text in Historical Document Recognition
di: Do, Thao, et al.
Pubblicazione: (2024)
di: Do, Thao, et al.
Pubblicazione: (2024)
Intriguing Properties of Large Language and Vision Models
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
di: Lee, Young-Jun, et al.
Pubblicazione: (2024)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
di: Liang, Qiao, et al.
Pubblicazione: (2025)
di: Liang, Qiao, et al.
Pubblicazione: (2025)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Bayesian Principles Improve Prompt Learning In Vision-Language Models
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
di: Kim, Mingyu, et al.
Pubblicazione: (2025)
LightHCG: a Lightweight yet powerful HSIC Disentanglement based Causal Glaucoma Detection Model framework
di: Kim, Daeyoung
Pubblicazione: (2025)
di: Kim, Daeyoung
Pubblicazione: (2025)
Do Vision-Language Models Understand Visual Persuasiveness?
di: Park, Gyuwon
Pubblicazione: (2025)
di: Park, Gyuwon
Pubblicazione: (2025)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
di: Kim, Dain, et al.
Pubblicazione: (2026)
di: Kim, Dain, et al.
Pubblicazione: (2026)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
di: Hwang, Taebaek, et al.
Pubblicazione: (2025)
di: Hwang, Taebaek, et al.
Pubblicazione: (2025)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
GCVAMD: A Modified CausalVAE Model for Causal Age-related Macular Degeneration Risk Factor Detection and Prediction
di: Kim, Daeyoung
Pubblicazione: (2025)
di: Kim, Daeyoung
Pubblicazione: (2025)
Vision-Language Models Do Not Understand Negation
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
di: Kim, Jungeun, et al.
Pubblicazione: (2024)
di: Kim, Jungeun, et al.
Pubblicazione: (2024)
LungCRCT: Causal Representation based Lung CT Processing for Lung Cancer Treatment
di: Kim, Daeyoung
Pubblicazione: (2026)
di: Kim, Daeyoung
Pubblicazione: (2026)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
di: Kim, Geewook, et al.
Pubblicazione: (2024)
di: Kim, Geewook, et al.
Pubblicazione: (2024)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
TroL: Traversal of Layers for Large Language and Vision Models
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
di: Lee, Byung-Kwan, et al.
Pubblicazione: (2024)
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
di: Gröpl, Marcel, et al.
Pubblicazione: (2026)
di: Gröpl, Marcel, et al.
Pubblicazione: (2026)
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
di: Kim, Hee-Seon, et al.
Pubblicazione: (2024)
di: Kim, Hee-Seon, et al.
Pubblicazione: (2024)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
di: Lee, Kang-il, et al.
Pubblicazione: (2024)
Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
di: Cho, Yeongjae, et al.
Pubblicazione: (2024)
VOMTC: Vision Objects for Millimeter and Terahertz Communications
di: Kim, Sunwoo, et al.
Pubblicazione: (2024)
di: Kim, Sunwoo, et al.
Pubblicazione: (2024)
Can Vision-Language Models Solve the Shell Game?
di: Liu, Tiedong, et al.
Pubblicazione: (2026)
di: Liu, Tiedong, et al.
Pubblicazione: (2026)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
di: Ghosh, Akash, et al.
Pubblicazione: (2024)
di: Ghosh, Akash, et al.
Pubblicazione: (2024)
CANVAS: A Benchmark for Vision-Language Models on Tool-Based User Interface Design
di: Jeong, Daeheon, et al.
Pubblicazione: (2025)
di: Jeong, Daeheon, et al.
Pubblicazione: (2025)
Focus Matters: Phase-Aware Suppression for Hallucination in Vision-Language Models
di: Kim, Sohyeon, et al.
Pubblicazione: (2026)
di: Kim, Sohyeon, et al.
Pubblicazione: (2026)
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
di: Park, Kwanyong, et al.
Pubblicazione: (2024)
di: Park, Kwanyong, et al.
Pubblicazione: (2024)
KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination
di: Choi, Byungjin, et al.
Pubblicazione: (2026)
di: Choi, Byungjin, et al.
Pubblicazione: (2026)
PROGRESSLM: Towards Progress Reasoning in Vision-Language Models
di: Zhang, Jianshu, et al.
Pubblicazione: (2026)
di: Zhang, Jianshu, et al.
Pubblicazione: (2026)
Vision Language Models are Biased
di: Vo, An, et al.
Pubblicazione: (2025)
di: Vo, An, et al.
Pubblicazione: (2025)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
di: Gao, Jianjun, et al.
Pubblicazione: (2024)
di: Gao, Jianjun, et al.
Pubblicazione: (2024)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
di: Park, Jonggwon, et al.
Pubblicazione: (2025)
di: Park, Jonggwon, et al.
Pubblicazione: (2025)
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
di: Kim, Junho, et al.
Pubblicazione: (2024)
di: Kim, Junho, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VARCO-VISION-2.0 Technical Report
di: Cha, Young-rok, et al.
Pubblicazione: (2025) -
V-Agent: An Interactive Video Search System Using Vision-Language Models
di: Park, SunYoung, et al.
Pubblicazione: (2025) -
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
di: Kim, Yoonshik, et al.
Pubblicazione: (2025) -
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
di: Lee, Seongyun, et al.
Pubblicazione: (2024) -
Evaluating Multimodal Generative AI with Korean Educational Standards
di: Park, Sanghee, et al.
Pubblicazione: (2025)