Gespeichert in:
| Hauptverfasser: | Wang, Yingzhi, Alhmoud, Anas, Alqurishi, Muhammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2412.13788 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
von: Wang, Yingzhi, et al.
Veröffentlicht: (2025)
von: Wang, Yingzhi, et al.
Veröffentlicht: (2025)
CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks
von: Qian, Zhaozhi, et al.
Veröffentlicht: (2024)
von: Qian, Zhaozhi, et al.
Veröffentlicht: (2024)
ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models
von: Alhumud, Anas, et al.
Veröffentlicht: (2026)
von: Alhumud, Anas, et al.
Veröffentlicht: (2026)
Zero-Shot Context-Aware ASR for Diverse Arabic Varieties
von: Talafha, Bashar, et al.
Veröffentlicht: (2025)
von: Talafha, Bashar, et al.
Veröffentlicht: (2025)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
von: Li, Haonan, et al.
Veröffentlicht: (2024)
von: Li, Haonan, et al.
Veröffentlicht: (2024)
Prompt-to-Leaderboard
von: Frick, Evan, et al.
Veröffentlicht: (2025)
von: Frick, Evan, et al.
Veröffentlicht: (2025)
The Leaderboard Illusion
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
von: Singh, Shivalika, et al.
Veröffentlicht: (2025)
CLARIN-PT-LDB: An Open LLM Leaderboard for Portuguese to assess Language, Culture and Civility
von: Silva, João, et al.
Veröffentlicht: (2026)
von: Silva, João, et al.
Veröffentlicht: (2026)
Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS
von: Al-Sabbagh, Rania
Veröffentlicht: (2026)
von: Al-Sabbagh, Rania
Veröffentlicht: (2026)
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
von: Grandury, María, et al.
Veröffentlicht: (2025)
von: Grandury, María, et al.
Veröffentlicht: (2025)
LEGOBench: Scientific Leaderboard Generation Benchmark
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
von: Abdoli, Sajjad, et al.
Veröffentlicht: (2026)
von: Abdoli, Sajjad, et al.
Veröffentlicht: (2026)
League: Leaderboard Generation on Demand
von: Wu, Jian, et al.
Veröffentlicht: (2025)
von: Wu, Jian, et al.
Veröffentlicht: (2025)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
Exploring the Latest LLMs for Leaderboard Extraction
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
von: Rao, Varun, et al.
Veröffentlicht: (2025)
von: Rao, Varun, et al.
Veröffentlicht: (2025)
User-centric Subjective Leaderboard by Customizable Reward Modeling
von: Jia, Qi, et al.
Veröffentlicht: (2025)
von: Jia, Qi, et al.
Veröffentlicht: (2025)
Improving LLM Leaderboards with Psychometrical Methodology
von: Federiakin, Denis
Veröffentlicht: (2025)
von: Federiakin, Denis
Veröffentlicht: (2025)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
von: Özyilmaz, Ömer Tarik, et al.
Veröffentlicht: (2025)
von: Özyilmaz, Ömer Tarik, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
Qabas: An Open-Source Arabic Lexicographic Database
von: Jarrar, Mustafa, et al.
Veröffentlicht: (2024)
von: Jarrar, Mustafa, et al.
Veröffentlicht: (2024)
Instruction Finetuning for Leaderboard Generation from Empirical AI Research
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards
von: Arzt, Varvara, et al.
Veröffentlicht: (2024)
von: Arzt, Varvara, et al.
Veröffentlicht: (2024)
MULTI: Multimodal Understanding Leaderboard with Text and Images
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
von: Zhu, Zichen, et al.
Veröffentlicht: (2024)
LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries
von: Wu, Zekun, et al.
Veröffentlicht: (2025)
von: Wu, Zekun, et al.
Veröffentlicht: (2025)
Creating Arabic LLM Prompts at Scale
von: El-Sheikh, Abdelrahman, et al.
Veröffentlicht: (2024)
von: El-Sheikh, Abdelrahman, et al.
Veröffentlicht: (2024)
The Trust Paradox: How CS Researchers Engage LLM Leaderboards
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2026)
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2026)
A Position Paper on the Automatic Generation of Machine Learning Leaderboards
von: Timmer, Roelien C, et al.
Veröffentlicht: (2025)
von: Timmer, Roelien C, et al.
Veröffentlicht: (2025)
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
von: Omnilingual ASR team, et al.
Veröffentlicht: (2025)
von: Omnilingual ASR team, et al.
Veröffentlicht: (2025)
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
Effective Context Selection in LLM-based Leaderboard Generation: An Empirical Study
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
CTC-DID: CTC-Based Arabic dialect identification for streaming applications
von: Farooq, Muhammad Umar, et al.
Veröffentlicht: (2026)
von: Farooq, Muhammad Umar, et al.
Veröffentlicht: (2026)
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
von: Jacovi, Alon, et al.
Veröffentlicht: (2025)
von: Jacovi, Alon, et al.
Veröffentlicht: (2025)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
von: Wang, Yingzhi, et al.
Veröffentlicht: (2025) -
CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks
von: Qian, Zhaozhi, et al.
Veröffentlicht: (2024) -
ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models
von: Alhumud, Anas, et al.
Veröffentlicht: (2026) -
Zero-Shot Context-Aware ASR for Diverse Arabic Varieties
von: Talafha, Bashar, et al.
Veröffentlicht: (2025) -
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
von: Srivastav, Vaibhav, et al.
Veröffentlicht: (2025)