Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Choi, Dasol, Son, Guijin, Kim, Soo Yong, Paik, Gio, Hong, Seunghyeok |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
Multi-Step Reasoning in Korean and the Emergent Mirage
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
di: Lee, Hanwool, et al.
Pubblicazione: (2025)
di: Lee, Hanwool, et al.
Pubblicazione: (2025)
MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
di: Paik, Gio, et al.
Pubblicazione: (2025)
di: Paik, Gio, et al.
Pubblicazione: (2025)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
HiKE: Hierarchical Evaluation Framework for Korean-English Code-Switching Speech Recognition
di: Paik, Gio, et al.
Pubblicazione: (2025)
di: Paik, Gio, et al.
Pubblicazione: (2025)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
Revisiting the UID Hypothesis in LLM Reasoning Traces
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
di: Choi, Dasol, et al.
Pubblicazione: (2026)
di: Choi, Dasol, et al.
Pubblicazione: (2026)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
di: Hong, Seokhee, et al.
Pubblicazione: (2025)
Visual Analytics for Fine-grained Text Classification Models and Datasets
di: Battogtokh, Munkhtulga, et al.
Pubblicazione: (2024)
di: Battogtokh, Munkhtulga, et al.
Pubblicazione: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
ESG Classification by Implicit Rule Learning via GPT-4
di: Yun, Hyo Jeong, et al.
Pubblicazione: (2024)
di: Yun, Hyo Jeong, et al.
Pubblicazione: (2024)
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
di: Tran, Hung Quang, et al.
Pubblicazione: (2026)
di: Tran, Hung Quang, et al.
Pubblicazione: (2026)
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Fine-grained Controllable Text Generation through In-context Learning with Feedback
di: Thillainathan, Sarubi, et al.
Pubblicazione: (2024)
di: Thillainathan, Sarubi, et al.
Pubblicazione: (2024)
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
Training Language Models to Generate Text with Citations via Fine-grained Rewards
di: Huang, Chengyu, et al.
Pubblicazione: (2024)
di: Huang, Chengyu, et al.
Pubblicazione: (2024)
Won: Establishing Best Practices for Korean Financial NLP
di: Son, Guijin, et al.
Pubblicazione: (2025)
di: Son, Guijin, et al.
Pubblicazione: (2025)
KAIO: A Collection of More Challenging Korean Questions
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
Controlling Language Confusion in Multilingual LLMs
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
di: Lee, Nahyun, et al.
Pubblicazione: (2025)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
di: Balauca, Ada-Astrid, et al.
Pubblicazione: (2024)
di: Balauca, Ada-Astrid, et al.
Pubblicazione: (2024)
No Language Data Left Behind: A Comparative Study of CJK Language Datasets in the Hugging Face Ecosystem
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
di: Shen, Yifan, et al.
Pubblicazione: (2025)
di: Shen, Yifan, et al.
Pubblicazione: (2025)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
di: Lee, Nahyun, et al.
Pubblicazione: (2026)
For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
di: Deng, Wenlong, et al.
Pubblicazione: (2025)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
di: Lee, Kyungjae, et al.
Pubblicazione: (2024)
di: Lee, Kyungjae, et al.
Pubblicazione: (2024)
Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models
di: Downer, Gabriel, et al.
Pubblicazione: (2025)
di: Downer, Gabriel, et al.
Pubblicazione: (2025)
Understanding Fine-grained Distortions in Reports of Scientific Findings
di: Wührl, Amelie, et al.
Pubblicazione: (2024)
di: Wührl, Amelie, et al.
Pubblicazione: (2024)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
di: Wang, Dianyi, et al.
Pubblicazione: (2025)
di: Wang, Dianyi, et al.
Pubblicazione: (2025)
ChartLens: Fine-grained Visual Attribution in Charts
di: Suri, Manan, et al.
Pubblicazione: (2025)
di: Suri, Manan, et al.
Pubblicazione: (2025)
SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL
di: Lee, Jimin, et al.
Pubblicazione: (2025)
di: Lee, Jimin, et al.
Pubblicazione: (2025)
Beyond Sentiment Classification: A Generative Framework for Emotion Intensity Evaluation in Text
di: Fabozzi, Francesco A., et al.
Pubblicazione: (2026)
di: Fabozzi, Francesco A., et al.
Pubblicazione: (2026)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
di: Ye, Junyi, et al.
Pubblicazione: (2024)
di: Ye, Junyi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
di: Ko, Hyunwoo, et al.
Pubblicazione: (2025) -
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
di: Choi, Dasol, et al.
Pubblicazione: (2025) -
Multi-Step Reasoning in Korean and the Emergent Mirage
di: Son, Guijin, et al.
Pubblicazione: (2025) -
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
di: Son, Guijin, et al.
Pubblicazione: (2024) -
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
di: Lee, Hanwool, et al.
Pubblicazione: (2025)