JEEM: Vision-Language Understanding in Four Arabic Dialects
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kadaoui, Karima, Atwany, Hanin, Al-Ali, Hamdan, Mohamed, Abdelrahman, Mekky, Ali, Tilga, Sergei, Fedorova, Natalia, Artemova, Ekaterina, Aldarmaki, Hanan, Kementchedjhieva, Yova |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
par: Ivanova, Anastasiia, et autres
Publié: (2025)
par: Ivanova, Anastasiia, et autres
Publié: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
par: Mohamed, Abdelrahman, et autres
Publié: (2025)
par: Mohamed, Abdelrahman, et autres
Publié: (2025)
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
par: Artemova, Ekaterina, et autres
Publié: (2024)
par: Artemova, Ekaterina, et autres
Publié: (2024)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
par: Agro, Maha Tufail, et autres
Publié: (2025)
par: Agro, Maha Tufail, et autres
Publié: (2025)
Dialectal Coverage And Generalization in Arabic Speech Recognition
par: Djanibekov, Amirbek, et autres
Publié: (2024)
par: Djanibekov, Amirbek, et autres
Publié: (2024)
ArFake: A Multi-Dialect Benchmark and Baselines for Arabic Spoof-Speech Detection
par: Maged, Mohamed, et autres
Publié: (2025)
par: Maged, Mohamed, et autres
Publié: (2025)
Mixat: A Data Set of Bilingual Emirati-English Speech
par: Ali, Maryam Al, et autres
Publié: (2024)
par: Ali, Maryam Al, et autres
Publié: (2024)
Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models
par: Mekky, Ali, et autres
Publié: (2026)
par: Mekky, Ali, et autres
Publié: (2026)
LLMs Can Compensate for Deficiencies in Visual Representations
par: Takishita, Sho, et autres
Publié: (2025)
par: Takishita, Sho, et autres
Publié: (2025)
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
par: Shoer, Belal, et autres
Publié: (2025)
par: Shoer, Belal, et autres
Publié: (2025)
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
par: Huzaifa, Muhammad, et autres
Publié: (2024)
par: Huzaifa, Muhammad, et autres
Publié: (2024)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
par: Salazar, Israfel, et autres
Publié: (2025)
par: Salazar, Israfel, et autres
Publié: (2025)
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
par: Toyin, Hawau Olamide, et autres
Publié: (2025)
par: Toyin, Hawau Olamide, et autres
Publié: (2025)
Personal Attribute Leakage in Federated Speech Models
par: Al-Ali, Hamdan, et autres
Publié: (2025)
par: Al-Ali, Hamdan, et autres
Publié: (2025)
CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
par: Ibrahim, George, et autres
Publié: (2025)
par: Ibrahim, George, et autres
Publié: (2025)
Answerability in Retrieval-Augmented Open-Domain Question Answering
par: Abdumalikov, Rustam, et autres
Publié: (2024)
par: Abdumalikov, Rustam, et autres
Publié: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
par: Chen, Xiaofu, et autres
Publié: (2025)
par: Chen, Xiaofu, et autres
Publié: (2025)
Realising metamorphic transformation in the mirror of Tamkeen: Growing a shared understanding from co‐reflected lived experiences
par: Louis Klein, et autres
Publié: (2024)
par: Louis Klein, et autres
Publié: (2024)
Beemo: Benchmark of Expert-edited Machine-generated Outputs
par: Artemova, Ekaterina, et autres
Publié: (2024)
par: Artemova, Ekaterina, et autres
Publié: (2024)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
par: Chernyshev, Konstantin, et autres
Publié: (2024)
par: Chernyshev, Konstantin, et autres
Publié: (2024)
Tendem: A Hybrid AI+Human Platform
par: Chernyshev, Konstantin, et autres
Publié: (2026)
par: Chernyshev, Konstantin, et autres
Publié: (2026)
Linear Semantic Segmentation for Low-Resource Spoken Dialects
par: Chirkunov, Kirill, et autres
Publié: (2026)
par: Chirkunov, Kirill, et autres
Publié: (2026)
Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs
par: Alakeel, Yara, et autres
Publié: (2026)
par: Alakeel, Yara, et autres
Publié: (2026)
On the Robust Approximation of ASR Metrics
par: Waheed, Abdul, et autres
Publié: (2025)
par: Waheed, Abdul, et autres
Publié: (2025)
What Do Speech Foundation Models Not Learn About Speech?
par: Waheed, Abdul, et autres
Publié: (2024)
par: Waheed, Abdul, et autres
Publié: (2024)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
par: Li, Jiaang, et autres
Publié: (2023)
par: Li, Jiaang, et autres
Publié: (2023)
Multimodal Large Language Models to Support Real-World Fact-Checking
par: Geng, Jiahui, et autres
Publié: (2024)
par: Geng, Jiahui, et autres
Publié: (2024)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
par: Ruder, Sebastian, et autres
Publié: (2018)
par: Ruder, Sebastian, et autres
Publié: (2018)
SparQLe: Speech Queries to Text Translation Through LLMs
par: Djanibekov, Amirbek, et autres
Publié: (2025)
par: Djanibekov, Amirbek, et autres
Publié: (2025)
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
par: Aldarmaki, Ibrahim, et autres
Publié: (2024)
par: Aldarmaki, Ibrahim, et autres
Publié: (2024)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
par: Sadallah, Abdelrahman, et autres
Publié: (2026)
par: Sadallah, Abdelrahman, et autres
Publié: (2026)
Overcoming Vocabulary Constraints with Pixel-level Fallback
par: Lotz, Jonas F., et autres
Publié: (2025)
par: Lotz, Jonas F., et autres
Publié: (2025)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
par: Toyin, Hawau Olamide, et autres
Publié: (2025)
par: Toyin, Hawau Olamide, et autres
Publié: (2025)
To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
par: Waheed, Abdul, et autres
Publié: (2024)
par: Waheed, Abdul, et autres
Publié: (2024)
Spoken Word2Vec: Learning Skipgram Embeddings from Speech
par: Sayeed, Mohammad Amaan, et autres
Publié: (2023)
par: Sayeed, Mohammad Amaan, et autres
Publié: (2023)
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models
par: Atwany, Hanin, et autres
Publié: (2025)
par: Atwany, Hanin, et autres
Publié: (2025)
MuLan: A Study of Fact Mutability in Language Models
par: Fierro, Constanza, et autres
Publié: (2024)
par: Fierro, Constanza, et autres
Publié: (2024)
Automatic Restoration of Diacritics for Speech Data Sets
par: Shatnawi, Sara, et autres
Publié: (2023)
par: Shatnawi, Sara, et autres
Publié: (2023)
LUNA: A Framework for Language Understanding and Naturalness Assessment
par: Saidov, Marat, et autres
Publié: (2024)
par: Saidov, Marat, et autres
Publié: (2024)
Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data
par: Peng, Siyao, et autres
Publié: (2024)
par: Peng, Siyao, et autres
Publié: (2024)
Documents similaires
-
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
par: Ivanova, Anastasiia, et autres
Publié: (2025) -
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
par: Mohamed, Abdelrahman, et autres
Publié: (2025) -
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
par: Artemova, Ekaterina, et autres
Publié: (2024) -
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
par: Agro, Maha Tufail, et autres
Publié: (2025) -
Dialectal Coverage And Generalization in Arabic Speech Recognition
par: Djanibekov, Amirbek, et autres
Publié: (2024)