Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Ersoy, Asim, Altinisik, Enes, Sencar, Husrev Taha, Darwish, Kareem |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
by: Fatehkia, Masoomali, et al.
Published: (2025)
by: Fatehkia, Masoomali, et al.
Published: (2025)
PAM: Training Policy-Aligned Moderation Filters at Scale
by: Fatehkia, Masoomali, et al.
Published: (2025)
by: Fatehkia, Masoomali, et al.
Published: (2025)
Explaining the role of Intrinsic Dimensionality in Adversarial Training
by: Altinisik, Enes, et al.
Published: (2024)
by: Altinisik, Enes, et al.
Published: (2024)
More Data, Fewer Diacritics: Scaling Arabic TTS
by: Musleh, Ahmed, et al.
Published: (2026)
by: Musleh, Ahmed, et al.
Published: (2026)
Multimedia Forensics
by: Husrev Taha Sencar
by: Husrev Taha Sencar
There Is More to Refusal in Large Language Models than a Single Direction
by: Joad, Faaiz, et al.
Published: (2026)
by: Joad, Faaiz, et al.
Published: (2026)
Organic Data-Driven Approach for Turkish Grammatical Error Correction and LLMs
by: Ersoy, Asım, et al.
Published: (2024)
by: Ersoy, Asım, et al.
Published: (2024)
Fanar 2.0: Arabic Generative AI Stack
by: FANAR TEAM, et al.
Published: (2026)
by: FANAR TEAM, et al.
Published: (2026)
Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction
by: Altinisik, Enes, et al.
Published: (2026)
by: Altinisik, Enes, et al.
Published: (2026)
Creating Arabic LLM Prompts at Scale
by: El-Sheikh, Abdelrahman, et al.
Published: (2024)
by: El-Sheikh, Abdelrahman, et al.
Published: (2024)
GemmAr: Enhancing LLMs Through Arabic Instruction-Tuning
by: Chouikhi, Hasna, et al.
Published: (2024)
by: Chouikhi, Hasna, et al.
Published: (2024)
From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction
by: Lekssays, Ahmed, et al.
Published: (2025)
by: Lekssays, Ahmed, et al.
Published: (2025)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
by: Chen, Mingyang, et al.
Published: (2024)
by: Chen, Mingyang, et al.
Published: (2024)
Call for Rigor in Reporting Quality of Instruction Tuning Data
by: Moon, Hyeonseok, et al.
Published: (2025)
by: Moon, Hyeonseok, et al.
Published: (2025)
Fanar: An Arabic-Centric Multimodal Generative AI Platform
by: Fanar Team, et al.
Published: (2025)
by: Fanar Team, et al.
Published: (2025)
Instruction-Guided Poetry Generation in Arabic and Its Dialects
by: Sadallah, Abdelrahman, et al.
Published: (2026)
by: Sadallah, Abdelrahman, et al.
Published: (2026)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
by: Bhatti, Hunzalah Hassan, et al.
Published: (2026)
by: Bhatti, Hunzalah Hassan, et al.
Published: (2026)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
by: Elozeiri, Kareem, et al.
Published: (2025)
by: Elozeiri, Kareem, et al.
Published: (2025)
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
by: Chen, Yi-Chang, et al.
Published: (2024)
by: Chen, Yi-Chang, et al.
Published: (2024)
When2Call: When (not) to Call Tools
by: Ross, Hayley, et al.
Published: (2025)
by: Ross, Hayley, et al.
Published: (2025)
SMART: Submodular Data Mixture Strategy for Instruction Tuning
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2024)
Instruction-Tuning LLMs for Event Extraction with Annotation Guidelines
by: Srivastava, Saurabh, et al.
Published: (2025)
by: Srivastava, Saurabh, et al.
Published: (2025)
Does Instruction Tuning Make LLMs More Consistent?
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Gazelle: An Instruction Dataset for Arabic Writing Assistance
by: Magdy, Samar M., et al.
Published: (2024)
by: Magdy, Samar M., et al.
Published: (2024)
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment
by: Elmadani, Khalid N., et al.
Published: (2025)
by: Elmadani, Khalid N., et al.
Published: (2025)
Strategies for Arabic Readability Modeling
by: Liberato, Juan Piñeros, et al.
Published: (2024)
by: Liberato, Juan Piñeros, et al.
Published: (2024)
TechniqueRAG: Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
by: Lekssays, Ahmed, et al.
Published: (2025)
by: Lekssays, Ahmed, et al.
Published: (2025)
The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
by: Lim, Henry, et al.
Published: (2025)
by: Lim, Henry, et al.
Published: (2025)
Balancing Continuous Pre-Training and Instruction Fine-Tuning: Optimizing Instruction-Following in LLMs
by: Jindal, Ishan, et al.
Published: (2024)
by: Jindal, Ishan, et al.
Published: (2024)
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
by: Le, Chenqian, et al.
Published: (2025)
by: Le, Chenqian, et al.
Published: (2025)
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai
by: Pengpun, Parinthapat, et al.
Published: (2024)
by: Pengpun, Parinthapat, et al.
Published: (2024)
Large-Scale Data Selection for Instruction Tuning
by: Ivison, Hamish, et al.
Published: (2025)
by: Ivison, Hamish, et al.
Published: (2025)
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation
by: Habash, Nizar, et al.
Published: (2024)
by: Habash, Nizar, et al.
Published: (2024)
Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications
by: Kanjirangat, Vani, et al.
Published: (2025)
by: Kanjirangat, Vani, et al.
Published: (2025)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
by: Crouse, Maxwell, et al.
Published: (2026)
by: Crouse, Maxwell, et al.
Published: (2026)
ArEEG_Chars: Dataset for Envisioned Speech Recognition using EEG for Arabic Characters
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
CIDAR: Culturally Relevant Instruction Dataset For Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
FinSphere, a Real-Time Stock Analysis Agent Powered by Instruction-Tuned LLMs and Domain Tools
by: Han, Shijie, et al.
Published: (2025)
by: Han, Shijie, et al.
Published: (2025)
Similar Items
-
FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
by: Fatehkia, Masoomali, et al.
Published: (2025) -
PAM: Training Policy-Aligned Moderation Filters at Scale
by: Fatehkia, Masoomali, et al.
Published: (2025) -
Explaining the role of Intrinsic Dimensionality in Adversarial Training
by: Altinisik, Enes, et al.
Published: (2024) -
More Data, Fewer Diacritics: Scaling Arabic TTS
by: Musleh, Ahmed, et al.
Published: (2026) -
Multimedia Forensics
by: Husrev Taha Sencar