RiddleBench: A New Generative Reasoning Benchmark for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Halder, Deepon, Saji, Alan, Jayakumar, Thanmay, Puduppully, Ratish, Kunchukuttan, Anoop, Dabre, Raj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026)
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026)
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
von: Halder, Deepon, et al.
Veröffentlicht: (2025)
Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026)
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026)
RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via Romanization
von: Husain, Jaavid Aktar, et al.
Veröffentlicht: (2024)
von: Husain, Jaavid Aktar, et al.
Veröffentlicht: (2024)
Top-b: Entropic Regulation of Relative Probability Bands in Autoregressive Language Processes
von: Halder, Deepon, et al.
Veröffentlicht: (2026)
von: Halder, Deepon, et al.
Veröffentlicht: (2026)
How Good is Zero-Shot MT Evaluation for Low Resource Indian Languages?
von: Singh, Anushka, et al.
Veröffentlicht: (2024)
von: Singh, Anushka, et al.
Veröffentlicht: (2024)
An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models
von: Mundra, Nandini, et al.
Veröffentlicht: (2024)
von: Mundra, Nandini, et al.
Veröffentlicht: (2024)
Airavata: Introducing Hindi Instruction-tuned LLM
von: Gala, Jay, et al.
Veröffentlicht: (2024)
von: Gala, Jay, et al.
Veröffentlicht: (2024)
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
von: Prasanjith, Pasunuti, et al.
Veröffentlicht: (2025)
von: Prasanjith, Pasunuti, et al.
Veröffentlicht: (2025)
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
von: Doddapaneni, Sumanth, et al.
Veröffentlicht: (2024)
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
von: Suryanarayanan, Sanjay, et al.
Veröffentlicht: (2024)
von: Suryanarayanan, Sanjay, et al.
Veröffentlicht: (2024)
Multilingual TinyStories: A Synthetic Combinatorial Corpus of Indic Children's Stories for Training Small Language Models
von: Halder, Deepon, et al.
Veröffentlicht: (2026)
von: Halder, Deepon, et al.
Veröffentlicht: (2026)
Leveraging Linguistically Enhanced Embeddings for Open Information Extraction
von: Farooqui, Fauzan, et al.
Veröffentlicht: (2024)
von: Farooqui, Fauzan, et al.
Veröffentlicht: (2024)
The Riddle of Reflection: Evaluating Reasoning and Self-Awareness in Multilingual LLMs using Indian Riddles
von: M, Abhinav P, et al.
Veröffentlicht: (2025)
von: M, Abhinav P, et al.
Veröffentlicht: (2025)
An Empirical Study of In-context Learning in LLMs for Machine Translation
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2024)
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2024)
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages
von: Sankar, Ashwin, et al.
Veröffentlicht: (2024)
von: Sankar, Ashwin, et al.
Veröffentlicht: (2024)
The Illusion of Generalization in Tabular Language Models
von: Gorla, Aditya, et al.
Veröffentlicht: (2026)
von: Gorla, Aditya, et al.
Veröffentlicht: (2026)
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities
von: Sravanthi, Settaluri Lakshmi, et al.
Veröffentlicht: (2024)
von: Sravanthi, Settaluri Lakshmi, et al.
Veröffentlicht: (2024)
NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
von: Boateng, George, et al.
Veröffentlicht: (2026)
von: Boateng, George, et al.
Veröffentlicht: (2026)
Can LLMs Solve My Grandma's Riddle? Evaluating Multilingual Large Language Models on Reasoning Traditional Bangla Tricky Riddles
von: Sayeedi, Nurul Labib, et al.
Veröffentlicht: (2025)
von: Sayeedi, Nurul Labib, et al.
Veröffentlicht: (2025)
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
von: Ghosh, Poulami, et al.
Veröffentlicht: (2024)
von: Ghosh, Poulami, et al.
Veröffentlicht: (2024)
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
von: Kartik, Kartik, et al.
Veröffentlicht: (2024)
von: Kartik, Kartik, et al.
Veröffentlicht: (2024)
CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2023)
von: Maurya, Kaushal Kumar, et al.
Veröffentlicht: (2023)
Pretraining Language Models Using Translationese
von: Doshi, Meet, et al.
Veröffentlicht: (2024)
von: Doshi, Meet, et al.
Veröffentlicht: (2024)
IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2024)
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2024)
Riddle Generation using Learning Resources
von: Parasa, Niharika Sri, et al.
Veröffentlicht: (2023)
von: Parasa, Niharika Sri, et al.
Veröffentlicht: (2023)
A Morphology-Based Investigation of Positional Encodings
von: Ghosh, Poulami, et al.
Veröffentlicht: (2024)
von: Ghosh, Poulami, et al.
Veröffentlicht: (2024)
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2025)
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2026)
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
von: Le, Duy, et al.
Veröffentlicht: (2025)
von: Le, Duy, et al.
Veröffentlicht: (2025)
How effective is Multi-source pivoting for Translation of Low Resource Indian Languages?
von: Gaikwad, Pranav, et al.
Veröffentlicht: (2024)
von: Gaikwad, Pranav, et al.
Veröffentlicht: (2024)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
BenchBench: Benchmarking Automated Benchmark Generation
von: Zheng, Yandan, et al.
Veröffentlicht: (2026)
von: Zheng, Yandan, et al.
Veröffentlicht: (2026)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
von: Gao, Zihan, et al.
Veröffentlicht: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI
von: Saji, Alan, et al.
Veröffentlicht: (2025) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025) -
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026) -
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
von: Halder, Deepon, et al.
Veröffentlicht: (2025) -
Scripts Through Time: A Survey of the Evolving Role of Transliteration in NLP
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026)