DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Khoramfar, Ali, Ramezani, Ali, Mohajeri, Mohammad Mahdi, Dousti, Mohammad Javad, Ahmadabadi, Majid Nili, Faili, Heshaam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoCoP: Enhancing Text Classification with LLM through Code Completion Prompt
by: Mohajeri, Mohammad Mahdi, et al.
Published: (2024)
by: Mohajeri, Mohammad Mahdi, et al.
Published: (2024)
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation
by: Seraj, Javad, et al.
Published: (2024)
by: Seraj, Javad, et al.
Published: (2024)
Mismatching-Aware Unsupervised Translation Quality Estimation For Low-Resource Languages
by: Azadi, Fatemeh, et al.
Published: (2022)
by: Azadi, Fatemeh, et al.
Published: (2022)
Persian Typographical Error Type Detection Using Deep Neural Networks on Algorithmically-Generated Misspellings
by: Dehghani, Mohammad, et al.
Published: (2023)
by: Dehghani, Mohammad, et al.
Published: (2023)
A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
by: Salehoof, Amir Mohammad, et al.
Published: (2025)
by: Salehoof, Amir Mohammad, et al.
Published: (2025)
$D^2LoRA$: Data-Driven LoRA Initialization for Low Resource Tasks
by: SeraJ, Javad, et al.
Published: (2025)
by: SeraJ, Javad, et al.
Published: (2025)
PersianPunc: A Large-Scale Dataset and BERT-Based Approach for Persian Punctuation Restoration
by: Kalahroodi, Mohammad Javad Ranjbar, et al.
Published: (2026)
by: Kalahroodi, Mohammad Javad Ranjbar, et al.
Published: (2026)
AI-powered Digital Framework for Personalized Economical Quality Learning at Scale
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark
by: Kalahroodi, Mohammad Javad Ranjbar, et al.
Published: (2026)
by: Kalahroodi, Mohammad Javad Ranjbar, et al.
Published: (2026)
PerSHOP -- A Persian dataset for shopping dialogue systems modeling
by: Mahmoudi, Keyvan, et al.
Published: (2024)
by: Mahmoudi, Keyvan, et al.
Published: (2024)
Three Benefits of Using Nonlinear Compliance in Robotic Systems Performing Cyclic Tasks: Energy Efficiency, Control Robustness, and Gait Optimality
by: Rezvan Nasiri, et al.
Published: (2025)
by: Rezvan Nasiri, et al.
Published: (2025)
PersianMind: A Cross-Lingual Persian-English Large Language Model
by: Rostami, Pedram, et al.
Published: (2024)
by: Rostami, Pedram, et al.
Published: (2024)
SearchInstruct: Enhancing Domain Adaptation via Retrieval-Based Instruction Dataset Creation
by: Barati, Iman, et al.
Published: (2025)
by: Barati, Iman, et al.
Published: (2025)
Dynamic Jointly Batch Selection for Data Efficient Machine Translation Fine-Tuning
by: Ghanizadeh, Mohammad Amin, et al.
Published: (2025)
by: Ghanizadeh, Mohammad Amin, et al.
Published: (2025)
Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning
by: Ghanizadeh, Mohammad Amin, et al.
Published: (2025)
by: Ghanizadeh, Mohammad Amin, et al.
Published: (2025)
CULL-MT: Compression Using Language and Layer pruning for Machine Translation
by: Rostami, Pedram, et al.
Published: (2024)
by: Rostami, Pedram, et al.
Published: (2024)
The use of the Extended Generalized Lambda Distribution for controlling the statistical process in individual measurements
by: Noorian, Sajad, et al.
Published: (2018)
by: Noorian, Sajad, et al.
Published: (2018)
Risk Sensitivity in Markov Games and Multi-Agent Reinforcement Learning: A Systematic Review
by: Ghaemi, Hafez, et al.
Published: (2024)
by: Ghaemi, Hafez, et al.
Published: (2024)
SchemaGraphSQL: Efficient Schema Linking with Pathfinding Graph Algorithms for Text-to-SQL on Large-Scale Databases
by: Safdarian, AmirHossein, et al.
Published: (2025)
by: Safdarian, AmirHossein, et al.
Published: (2025)
A Survey On Enhancing Reinforcement Learning in Complex Environments: Insights from Human and LLM Feedback
by: Laleh, Alireza Rashidi, et al.
Published: (2024)
by: Laleh, Alireza Rashidi, et al.
Published: (2024)
Matina: A Large-Scale 73B Token Persian Text Corpus
by: Hosseinbeigi, Sara Bourbour, et al.
Published: (2025)
by: Hosseinbeigi, Sara Bourbour, et al.
Published: (2025)
WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition
by: Ramezani, Erfan, et al.
Published: (2026)
by: Ramezani, Erfan, et al.
Published: (2026)
Airfoil Shape Optimization in Ultralow Reynolds Flows Applying a Deep Learning–Genetic Algorithm Framework on a Shear‐Stress‐Based Inverse Design Method
by: Zakaria Drafsh, et al.
Published: (2025)
by: Zakaria Drafsh, et al.
Published: (2025)
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
by: Moradi, Mohammad Mahdi, et al.
Published: (2025)
by: Moradi, Mohammad Mahdi, et al.
Published: (2025)
Deep Learning-based Sentiment Analysis in Persian Language
by: Heydari, Mohammad, et al.
Published: (2024)
by: Heydari, Mohammad, et al.
Published: (2024)
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
by: Mirbagheri, Mohammad Reza, et al.
Published: (2025)
by: Mirbagheri, Mohammad Reza, et al.
Published: (2025)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering
by: Astaraki, Mahdi, et al.
Published: (2026)
by: Astaraki, Mahdi, et al.
Published: (2026)
Fine-Tuning LLMs for Reliable Medical Question-Answering Services
by: Anaissi, Ali, et al.
Published: (2024)
by: Anaissi, Ali, et al.
Published: (2024)
Enhancing Few-Shot Transfer Learning with Optimized Multi-Task Prompt Tuning through Modular Prompt Composition
by: Pouramini, Ahmad, et al.
Published: (2024)
by: Pouramini, Ahmad, et al.
Published: (2024)
CrossPT: Exploring Cross-Task Transferability through Multi-Task Prompt Tuning
by: Pouramini, Ahmad, et al.
Published: (2025)
by: Pouramini, Ahmad, et al.
Published: (2025)
SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions
by: Liu, Shicheng, et al.
Published: (2024)
by: Liu, Shicheng, et al.
Published: (2024)
Towards Domain Specification of Embedding Models in Medicine
by: Khodadad, Mohammad, et al.
Published: (2025)
by: Khodadad, Mohammad, et al.
Published: (2025)
Evaluating the Creativity of LLMs in Persian Literary Text Generation
by: Tourajmehr, Armin, et al.
Published: (2025)
by: Tourajmehr, Armin, et al.
Published: (2025)
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language
by: Jamali, Naghmeh, et al.
Published: (2025)
by: Jamali, Naghmeh, et al.
Published: (2025)
Evaluating Multi-Hop Reasoning in Large Language Models: A Chemistry-Centric Case Study
by: Khodadad, Mohammad, et al.
Published: (2025)
by: Khodadad, Mohammad, et al.
Published: (2025)
Benchmarking ChatGPT and DeepSeek in April 2025: A Novel Dual Perspective Sentiment Analysis Using Lexicon-Based and Deep Learning Approaches
by: Alhusseini, Maryam Mahdi, et al.
Published: (2025)
by: Alhusseini, Maryam Mahdi, et al.
Published: (2025)
MultiMind at SemEval-2025 Task 7: Crosslingual Fact-Checked Claim Retrieval via Multi-Source Alignment
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2025)
Subgoal Discovery Using a Free Energy Paradigm and State Aggregations
by: Mesbah, Amirhossein, et al.
Published: (2024)
by: Mesbah, Amirhossein, et al.
Published: (2024)
Similar Items
-
CoCoP: Enhancing Text Classification with LLM through Code Completion Prompt
by: Mohajeri, Mohammad Mahdi, et al.
Published: (2024) -
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation
by: Seraj, Javad, et al.
Published: (2024) -
Mismatching-Aware Unsupervised Translation Quality Estimation For Low-Resource Languages
by: Azadi, Fatemeh, et al.
Published: (2022) -
Persian Typographical Error Type Detection Using Deep Neural Networks on Algorithmically-Generated Misspellings
by: Dehghani, Mohammad, et al.
Published: (2023) -
A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
by: Salehoof, Amir Mohammad, et al.
Published: (2025)