Saved in:
| Main Authors: | Alyafeai, Zaid, Pieler, Michael, Teufel, Hannah, Tow, Jonathan, Bellagente, Marco, Phung, Duy, Pinnaparaju, Nikhil, Adithyan, Reshinth, Rocha, Paulo, Zhuravinskyi, Maksym, Riquelme, Carlos |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.04277 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
by: Pieler, Michael, et al.
Published: (2024)
by: Pieler, Michael, et al.
Published: (2024)
Stable LM 2 1.6B Technical Report
by: Bellagente, Marco, et al.
Published: (2024)
by: Bellagente, Marco, et al.
Published: (2024)
Stable Code Technical Report
by: Pinnaparaju, Nikhil, et al.
Published: (2024)
by: Pinnaparaju, Nikhil, et al.
Published: (2024)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
by: Chatterjee, Agneet, et al.
Published: (2025)
by: Chatterjee, Agneet, et al.
Published: (2025)
3LM: Bridging Arabic, STEM, and Code through Benchmarking
by: Boussaha, Basma El Amel, et al.
Published: (2025)
by: Boussaha, Basma El Amel, et al.
Published: (2025)
Poem Meter Classification of Recited Arabic Poetry: Integrating High-Resource Systems for a Low-Resource Task
by: Al-Shaibani, Maged S., et al.
Published: (2025)
by: Al-Shaibani, Maged S., et al.
Published: (2025)
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
CIDAR: Culturally Relevant Instruction Dataset For Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
Adapting WavLM for Speech Emotion Recognition
by: Diatlova, Daria, et al.
Published: (2024)
by: Diatlova, Daria, et al.
Published: (2024)
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation
by: AlQadi, Leen, et al.
Published: (2026)
by: AlQadi, Leen, et al.
Published: (2026)
Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps
by: Alzubaidi, Ahmed, et al.
Published: (2025)
by: Alzubaidi, Ahmed, et al.
Published: (2025)
From Arabic Text to Puzzles: LLM-Driven Development of Arabic Educational Crosswords
by: Zeinalipour, Kamyar, et al.
Published: (2025)
by: Zeinalipour, Kamyar, et al.
Published: (2025)
MeXtract: Light-Weight Metadata Extraction from Scientific Papers
by: Alyafeai, Zaid, et al.
Published: (2025)
by: Alyafeai, Zaid, et al.
Published: (2025)
MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs
by: Alyafeai, Zaid, et al.
Published: (2025)
by: Alyafeai, Zaid, et al.
Published: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
by: Doan, Khai Duy, et al.
Published: (2024)
by: Doan, Khai Duy, et al.
Published: (2024)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates
by: Gu, Jian, et al.
Published: (2026)
by: Gu, Jian, et al.
Published: (2026)
Type-Compliant Adaptation Cascades: Adapting Programmatic LM Workflows to Data
by: Lin, Chu-Cheng, et al.
Published: (2025)
by: Lin, Chu-Cheng, et al.
Published: (2025)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
by: Jo, Daejin, et al.
Published: (2025)
by: Jo, Daejin, et al.
Published: (2025)
Lisan Al-Arab: Journal of Arabic Language and Arabic
Published: (2016)
Published: (2016)
Beyond Arabic: Software for Perso-Arabic Script Manipulation
by: Gutkin, Alexander, et al.
Published: (2023)
by: Gutkin, Alexander, et al.
Published: (2023)
AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
by: El-Haj, Mo, et al.
Published: (2025)
by: El-Haj, Mo, et al.
Published: (2025)
StyleAdaptedLM: Enhancing Instruction Following Models with Efficient Stylistic Transfer
by: Ramu, Pritika, et al.
Published: (2025)
by: Ramu, Pritika, et al.
Published: (2025)
BAMBINO-LM: (Bilingual-)Human-Inspired Continual Pretraining of BabyLM
by: Shen, Zhewen, et al.
Published: (2024)
by: Shen, Zhewen, et al.
Published: (2024)
Quranic Arabic
by: van Putten, Marijn
Published: (2022)
by: van Putten, Marijn
Published: (2022)
Arabic Dialogues
by: Mairs, Rachel
Published: (2024)
by: Mairs, Rachel
Published: (2024)
The Arabic Fable
by: Marzolph, Ulrich
Published: (2025)
by: Marzolph, Ulrich
Published: (2025)
Latin and Arabic
by: König, Daniel G.
Published: (2020)
by: König, Daniel G.
Published: (2020)
Arabic in Context
Published: (2025)
Published: (2025)
Arabic in Contact
Published: (2023)
Published: (2023)
Arabic Abstracts
Published: (2024)
Published: (2024)
Arabic Abstract
Published: (2024)
Published: (2024)
The Orientalists' Stance Towards Arabic Sciences (Especially Arabic Astronomy)
by: Abdullah, Duaa, et al.
Published: (2025)
by: Abdullah, Duaa, et al.
Published: (2025)
Arabic Little STT: Arabic Children Speech Recognition Dataset
by: Alkadri, Mouhand, et al.
Published: (2025)
by: Alkadri, Mouhand, et al.
Published: (2025)
The Arabic Generality Score: Another Dimension of Modeling Arabic Dialectness
by: Shaban, Sanad, et al.
Published: (2025)
by: Shaban, Sanad, et al.
Published: (2025)
Advancing Dialectal Arabic to Modern Standard Arabic Machine Translation
by: Alabdullah, Abdullah, et al.
Published: (2025)
by: Alabdullah, Abdullah, et al.
Published: (2025)
Replicating ReLM Results: Validating Large Language Models with ReLM
by: Adamson, Reece, et al.
Published: (2025)
by: Adamson, Reece, et al.
Published: (2025)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
by: Charpentier, Lucas, et al.
Published: (2025)
by: Charpentier, Lucas, et al.
Published: (2025)
BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs
by: Boizard, Nicolas, et al.
Published: (2026)
by: Boizard, Nicolas, et al.
Published: (2026)
Similar Items
-
Rephrasing natural text data with different languages and quality levels for Large Language Model pre-training
by: Pieler, Michael, et al.
Published: (2024) -
Stable LM 2 1.6B Technical Report
by: Bellagente, Marco, et al.
Published: (2024) -
Stable Code Technical Report
by: Pinnaparaju, Nikhil, et al.
Published: (2024) -
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
by: Chatterjee, Agneet, et al.
Published: (2025) -
3LM: Bridging Arabic, STEM, and Code through Benchmarking
by: Boussaha, Basma El Amel, et al.
Published: (2025)