Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Abdaljalil, Samir, Serpedin, Erchin, Qaraqe, Khalid, Kurban, Hasan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
by: Mahmood, Alhasan, et al.
Published: (2026)
by: Mahmood, Alhasan, et al.
Published: (2026)
4D Synchronized Fields: Motion-Language Gaussian Splatting for Temporal Scene Understanding
by: Barhdadi, Mohamed Rayan, et al.
Published: (2026)
by: Barhdadi, Mohamed Rayan, et al.
Published: (2026)
SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Stress-Testing Multimodal Foundation Models for Crystallographic Reasoning
by: Polat, Can, et al.
Published: (2025)
by: Polat, Can, et al.
Published: (2025)
QuantumCanvas: A Multimodal Benchmark for Visual Learning of Atomic Interactions
by: Polat, Can, et al.
Published: (2025)
by: Polat, Can, et al.
Published: (2025)
SCALAR: Quantifying Structural Hallucination, Consistency, and Reasoning Gaps in Materials Foundation Models
by: Polat, Can, et al.
Published: (2026)
by: Polat, Can, et al.
Published: (2026)
EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration
by: Barhdadi, Mohamed Rayan, et al.
Published: (2025)
by: Barhdadi, Mohamed Rayan, et al.
Published: (2025)
Code-Switching Curriculum Learning for Multilingual Transfer in LLMs
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
by: Yoo, Haneul, et al.
Published: (2024)
by: Yoo, Haneul, et al.
Published: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs
by: Guo, Yanzhu, et al.
Published: (2024)
by: Guo, Yanzhu, et al.
Published: (2024)
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
by: Koo, Hamin, et al.
Published: (2025)
by: Koo, Hamin, et al.
Published: (2025)
POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
by: Chen, Yizhuo, et al.
Published: (2025)
by: Chen, Yizhuo, et al.
Published: (2025)
xChemAgents: Agentic AI for Explainable Quantum Chemistry
by: Polat, Can, et al.
Published: (2025)
by: Polat, Can, et al.
Published: (2025)
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
by: Xie, Peng, et al.
Published: (2025)
by: Xie, Peng, et al.
Published: (2025)
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
by: Asano, Shunta, et al.
Published: (2026)
by: Asano, Shunta, et al.
Published: (2026)
Beyond Atomic Geometry Representations in Materials Science: A Human-in-the-Loop Multimodal Framework
by: Polat, Can, et al.
Published: (2025)
by: Polat, Can, et al.
Published: (2025)
Understanding the Capabilities of Molecular Graph Neural Networks in Materials Science Through Multimodal Learning and Physical Context Encoding
by: Polat, Can, et al.
Published: (2025)
by: Polat, Can, et al.
Published: (2025)
C2NP: A Benchmark for Learning Scale-Dependent Geometric Invariances in 3D Materials Generation
by: Polat, Can, et al.
Published: (2026)
by: Polat, Can, et al.
Published: (2026)
How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science
by: Polat, Can, et al.
Published: (2026)
by: Polat, Can, et al.
Published: (2026)
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
by: Zhang, Shimao, et al.
Published: (2025)
by: Zhang, Shimao, et al.
Published: (2025)
Conditioning LLMs to Generate Code-Switched Text
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation
by: Wu, Sophie, et al.
Published: (2026)
by: Wu, Sophie, et al.
Published: (2026)
SLAM: Towards Efficient Multilingual Reasoning via Selective Language Alignment
by: Fan, Yuchun, et al.
Published: (2025)
by: Fan, Yuchun, et al.
Published: (2025)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
by: Gupta, Ashray, et al.
Published: (2025)
by: Gupta, Ashray, et al.
Published: (2025)
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
by: Khanbayov, Rasul, et al.
Published: (2026)
by: Khanbayov, Rasul, et al.
Published: (2026)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
by: Katzy, Jonathan, et al.
Published: (2025)
by: Katzy, Jonathan, et al.
Published: (2025)
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
by: Han, HyoJung, et al.
Published: (2025)
by: Han, HyoJung, et al.
Published: (2025)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
by: Chehade, Mohamad, et al.
Published: (2025)
by: Chehade, Mohamad, et al.
Published: (2025)
Eliciting Better Multilingual Structured Reasoning from LLMs through Code
by: Li, Bryan, et al.
Published: (2024)
by: Li, Bryan, et al.
Published: (2024)
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
by: Gurgurov, Daniil, et al.
Published: (2024)
by: Gurgurov, Daniil, et al.
Published: (2024)
Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities
by: Gupta, Ayushman, et al.
Published: (2024)
by: Gupta, Ayushman, et al.
Published: (2024)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
by: Yan, Weixiang, et al.
Published: (2023)
by: Yan, Weixiang, et al.
Published: (2023)
Similar Items
-
Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
by: Abdaljalil, Samir, et al.
Published: (2025) -
Theorem-of-Thought: A Multi-Agent Framework for Abductive, Deductive, and Inductive Reasoning in Language Models
by: Abdaljalil, Samir, et al.
Published: (2025) -
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025) -
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
by: Abdaljalil, Samir, et al.
Published: (2026) -
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
by: Abdaljalil, Samir, et al.
Published: (2026)