Benchmarking Vision Language Models on German Factual Data
Fuente:
arXiv
Saved in:
| Main Authors: | Peinl, René, Tischler, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
by: Peinl, René, et al.
Published: (2025)
by: Peinl, René, et al.
Published: (2025)
JVLGS: Joint Vision-Language Gas Leak Segmentation
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
by: Ma, Yueen, et al.
Published: (2024)
by: Ma, Yueen, et al.
Published: (2024)
Single-image Reflectance and Transmittance Estimation from Any Flatbed Scanner
by: Rodriguez-Pardo, Carlos, et al.
Published: (2025)
by: Rodriguez-Pardo, Carlos, et al.
Published: (2025)
NeuBTF: Neural fields for BTF encoding and transfer
by: Rodriguez-Pardo, Carlos, et al.
Published: (2023)
by: Rodriguez-Pardo, Carlos, et al.
Published: (2023)
UMat: Uncertainty-Aware Single Image High Resolution Material Capture
by: Rodriguez-Pardo, Carlos, et al.
Published: (2023)
by: Rodriguez-Pardo, Carlos, et al.
Published: (2023)
Archival Faces: Detection of Faces in Digitized Historical Documents
by: Vaško, Marek, et al.
Published: (2025)
by: Vaško, Marek, et al.
Published: (2025)
Automatic infant 2D pose estimation from videos: comparing seven deep neural network methods
by: Gama, Filipe, et al.
Published: (2024)
by: Gama, Filipe, et al.
Published: (2024)
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
by: Louison, Nikita, et al.
Published: (2024)
by: Louison, Nikita, et al.
Published: (2024)
LRVS-Fashion: Extending Visual Search with Referring Instructions
by: Lepage, Simon, et al.
Published: (2023)
by: Lepage, Simon, et al.
Published: (2023)
VALE: A Multimodal Visual and Language Explanation Framework for Image Classifiers using eXplainable AI and Language Models
by: Natarajan, Purushothaman, et al.
Published: (2024)
by: Natarajan, Purushothaman, et al.
Published: (2024)
Handling Out-of-Distribution Data: A Survey
by: Tamang, Lakpa, et al.
Published: (2025)
by: Tamang, Lakpa, et al.
Published: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
by: Papyan, Narek, et al.
Published: (2024)
by: Papyan, Narek, et al.
Published: (2024)
Data Augmentation and Resolution Enhancement using GANs and Diffusion Models for Tree Segmentation
by: Ferreira, Alessandro dos Santos, et al.
Published: (2025)
by: Ferreira, Alessandro dos Santos, et al.
Published: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
Unlocking UML Class Diagram Understanding in Vision Language Models
by: Naboichenko, Artem, et al.
Published: (2026)
by: Naboichenko, Artem, et al.
Published: (2026)
Fine-grained spatial-temporal perception for gas leak segmentation
by: Zhao, Xinlong, et al.
Published: (2025)
by: Zhao, Xinlong, et al.
Published: (2025)
On Memory: A comparison of memory mechanisms in world models
by: Laird, Eli J., et al.
Published: (2025)
by: Laird, Eli J., et al.
Published: (2025)
HelloMeme: Integrating Spatial Knitting Attentions to Embed High-Level and Fidelity-Rich Conditions in Diffusion Models
by: Zhang, Shengkai, et al.
Published: (2024)
by: Zhang, Shengkai, et al.
Published: (2024)
Few-Class Arena: A Benchmark for Efficient Selection of Vision Models and Dataset Difficulty Measurement
by: Cao, Bryan Bo, et al.
Published: (2024)
by: Cao, Bryan Bo, et al.
Published: (2024)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
by: Chen, Kewei, et al.
Published: (2025)
by: Chen, Kewei, et al.
Published: (2025)
Image-based Facial Rig Inversion
by: Yang, Tianxiang, et al.
Published: (2025)
by: Yang, Tianxiang, et al.
Published: (2025)
Cost-Effective Attention Mechanisms for Low Resource Settings: Necessity & Sufficiency of Linear Transformations
by: Hosseini, Peyman, et al.
Published: (2024)
by: Hosseini, Peyman, et al.
Published: (2024)
Does CLIP perceive art the same way we do?
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
by: Malikussaid, et al.
Published: (2026)
by: Malikussaid, et al.
Published: (2026)
Robust Noise Attenuation via Adaptive Pooling of Transformer Outputs
by: Brothers, Greyson
Published: (2025)
by: Brothers, Greyson
Published: (2025)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
by: Kang, Xueyang, et al.
Published: (2026)
by: Kang, Xueyang, et al.
Published: (2026)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
by: Zhang, Junwen, et al.
Published: (2025)
by: Zhang, Junwen, et al.
Published: (2025)
Attention Gathers, MLPs Compose: A Causal Analysis of an Action-Outcome Circuit in VideoViT
by: Chereddy, Sai V R
Published: (2026)
by: Chereddy, Sai V R
Published: (2026)
FastGS: Training 3D Gaussian Splatting in 100 Seconds
by: Ren, Shiwei, et al.
Published: (2025)
by: Ren, Shiwei, et al.
Published: (2025)
SEGS-SLAM: Structure-enhanced 3D Gaussian Splatting SLAM with Appearance Embedding
by: Wen, Tianci, et al.
Published: (2025)
by: Wen, Tianci, et al.
Published: (2025)
Empowering Image Recovery_ A Multi-Attention Approach
by: Wen, Juan, et al.
Published: (2024)
by: Wen, Juan, et al.
Published: (2024)
Application of deep learning approaches for medieval historical documents transcription
by: Voloshchuk, Maksym, et al.
Published: (2025)
by: Voloshchuk, Maksym, et al.
Published: (2025)
A large-scale, physically-based synthetic dataset for satellite pose estimation
by: Velkei, Szabolcs, et al.
Published: (2025)
by: Velkei, Szabolcs, et al.
Published: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
by: Patel, Urjitkumar, et al.
Published: (2025)
by: Patel, Urjitkumar, et al.
Published: (2025)
Force-Aware 3D Contact Modeling for Stable Grasp Generation
by: Chen, Zhuo, et al.
Published: (2025)
by: Chen, Zhuo, et al.
Published: (2025)
Underwater SONAR Image Classification and Analysis using LIME-based Explainable Artificial Intelligence
by: Natarajan, Purushothaman, et al.
Published: (2024)
by: Natarajan, Purushothaman, et al.
Published: (2024)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
by: Li, Mingda, et al.
Published: (2024)
by: Li, Mingda, et al.
Published: (2024)
Multimodal Image Matching based on Frequency-domain Information of Local Energy Response
by: Yang, Meng, et al.
Published: (2025)
by: Yang, Meng, et al.
Published: (2025)
Similar Items
-
VLM@school -- Evaluation of AI image understanding on German middle school knowledge
by: Peinl, René, et al.
Published: (2025) -
JVLGS: Joint Vision-Language Gas Leak Segmentation
by: Zhao, Xinlong, et al.
Published: (2025) -
A Survey on Vision-Language-Action Models for Embodied AI
by: Ma, Yueen, et al.
Published: (2024) -
Single-image Reflectance and Transmittance Estimation from Any Flatbed Scanner
by: Rodriguez-Pardo, Carlos, et al.
Published: (2025) -
NeuBTF: Neural fields for BTF encoding and transfer
by: Rodriguez-Pardo, Carlos, et al.
Published: (2023)