Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen-Truong, Hai, Balbay, Alper, Bayrak, Tunga |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
VG3T: Visual Geometry Grounded Gaussian Transformer
von: Kim, Junho, et al.
Veröffentlicht: (2025)
von: Kim, Junho, et al.
Veröffentlicht: (2025)
XEdgeAI: A Human-centered Industrial Inspection Framework with Data-centric Explainable Edge AI Approach
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024)
Towards Understanding Visual Grounding in Visual Language Models
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2025)
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
von: Prasad, Archiki, et al.
Veröffentlicht: (2023)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
von: Lee, Byung-Kwan, et al.
Veröffentlicht: (2025)
Generative Visual Communication in the Era of Vision-Language Models
von: Vinker, Yael
Veröffentlicht: (2024)
von: Vinker, Yael
Veröffentlicht: (2024)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
von: Lagos, Maximiliano Hormazábal, et al.
Veröffentlicht: (2025)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
von: Quang, Ngoc Bui Lam, et al.
Veröffentlicht: (2025)
von: Quang, Ngoc Bui Lam, et al.
Veröffentlicht: (2025)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2025)
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2025)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models
von: Wang, Hongjun, et al.
Veröffentlicht: (2026)
von: Wang, Hongjun, et al.
Veröffentlicht: (2026)
Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models
von: Sanogo, Kassoum, et al.
Veröffentlicht: (2025)
von: Sanogo, Kassoum, et al.
Veröffentlicht: (2025)
Jailbreaking Vision-Language Models Through the Visual Modality
von: Azulay, Aharon, et al.
Veröffentlicht: (2026)
von: Azulay, Aharon, et al.
Veröffentlicht: (2026)
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
von: Li, Kevin, et al.
Veröffentlicht: (2025)
von: Li, Kevin, et al.
Veröffentlicht: (2025)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
von: Park, Jihwan, et al.
Veröffentlicht: (2025)
Cortex-Grounded Diffusion Models for Brain Image Generation
von: Bongratz, Fabian, et al.
Veröffentlicht: (2026)
von: Bongratz, Fabian, et al.
Veröffentlicht: (2026)
Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
von: Hossain, Shamima
Veröffentlicht: (2025)
von: Hossain, Shamima
Veröffentlicht: (2025)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
von: Cao, Shengcao, et al.
Veröffentlicht: (2024)
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
von: Liu, Zhongye, et al.
Veröffentlicht: (2024)
von: Liu, Zhongye, et al.
Veröffentlicht: (2024)
Towards Long-window Anchoring in Vision-Language Model Distillation
von: Zhou, Haoyi, et al.
Veröffentlicht: (2025)
von: Zhou, Haoyi, et al.
Veröffentlicht: (2025)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
von: Cui, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Cui, Kaiyuan, et al.
Veröffentlicht: (2026)
PhyGround: Benchmarking Physical Reasoning in Generative World Models
von: Lin, Juyi, et al.
Veröffentlicht: (2026)
von: Lin, Juyi, et al.
Veröffentlicht: (2026)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2025)
Towards Understanding and Quantifying Uncertainty for Text-to-Image Generation
von: Franchi, Gianni, et al.
Veröffentlicht: (2024)
von: Franchi, Gianni, et al.
Veröffentlicht: (2024)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
von: Tan, Zhaorui, et al.
Veröffentlicht: (2024)
Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
von: Güzel, Ahmet H., et al.
Veröffentlicht: (2025)
von: Güzel, Ahmet H., et al.
Veröffentlicht: (2025)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
von: Du, Zilin, et al.
Veröffentlicht: (2024)
von: Du, Zilin, et al.
Veröffentlicht: (2024)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
von: Sui, Elaine, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024) -
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025) -
VG3T: Visual Geometry Grounded Gaussian Transformer
von: Kim, Junho, et al.
Veröffentlicht: (2025) -
XEdgeAI: A Human-centered Industrial Inspection Framework with Data-centric Explainable Edge AI Approach
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2024) -
Towards Understanding Visual Grounding in Visual Language Models
von: Pantazopoulos, Georgios, et al.
Veröffentlicht: (2025)