NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors
Fuente:
arXiv
Saved in:
| Main Authors: | Naeem, Numaan, Ahmad, Sarfraz, Ahsan, Momina, Iqbal, Hasan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025)
by: Ahmad, Sarfraz, et al.
Published: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
by: Bouthors, Maxime, et al.
Published: (2025)
by: Bouthors, Maxime, et al.
Published: (2025)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
by: Nzeyimana, Antoine, et al.
Published: (2025)
by: Nzeyimana, Antoine, et al.
Published: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
by: Danilevsky, Marina, et al.
Published: (2025)
by: Danilevsky, Marina, et al.
Published: (2025)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
by: Dai, Xinbang, et al.
Published: (2025)
by: Dai, Xinbang, et al.
Published: (2025)
Adaptive Schema-aware Event Extraction with Retrieval-Augmented Generation
by: Liang, Sheng, et al.
Published: (2025)
by: Liang, Sheng, et al.
Published: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
POISONCRAFT: Practical Poisoning of Retrieval-Augmented Generation for Large Language Models
by: Shao, Yangguang, et al.
Published: (2025)
by: Shao, Yangguang, et al.
Published: (2025)
RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion
by: Niu, Guanglin, et al.
Published: (2026)
by: Niu, Guanglin, et al.
Published: (2026)
A Fuzzy Logic Prompting Framework for Large Language Models in Adaptive and Uncertain Tasks
by: Figueiredo, Vanessa
Published: (2025)
by: Figueiredo, Vanessa
Published: (2025)
Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
by: Kocbek, Primož, et al.
Published: (2025)
by: Kocbek, Primož, et al.
Published: (2025)
Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
by: Šindelář, Pavel, et al.
Published: (2025)
by: Šindelář, Pavel, et al.
Published: (2025)
Overview of the ClinIQLink 2025 Shared Task on Medical Question-Answering
by: Colelough, Brandon, et al.
Published: (2025)
by: Colelough, Brandon, et al.
Published: (2025)
Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
by: Riemenschneider, Frederick, et al.
Published: (2024)
by: Riemenschneider, Frederick, et al.
Published: (2024)
On the Limitations of Vision-Language Models in Understanding Image Transforms
by: Anis, Ahmad Mustafa, et al.
Published: (2025)
by: Anis, Ahmad Mustafa, et al.
Published: (2025)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
by: Wei, Xingyuan, et al.
Published: (2024)
by: Wei, Xingyuan, et al.
Published: (2024)
PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI
by: Chaitanya, Keshava, et al.
Published: (2026)
by: Chaitanya, Keshava, et al.
Published: (2026)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
by: Hildebrand, Samuel, et al.
Published: (2025)
by: Hildebrand, Samuel, et al.
Published: (2025)
Accurate and Energy Efficient: Local Retrieval-Augmented Generation Models Outperform Commercial Large Language Models in Medical Tasks
by: Vrettos, Konstantinos, et al.
Published: (2025)
by: Vrettos, Konstantinos, et al.
Published: (2025)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
by: Wang, Yongjie, et al.
Published: (2025)
by: Wang, Yongjie, et al.
Published: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Considering Length Diversity in Retrieval-Augmented Summarization
by: Juseon-Do, et al.
Published: (2025)
by: Juseon-Do, et al.
Published: (2025)
VERA: Validation and Evaluation of Retrieval-Augmented Systems
by: Ding, Tianyu, et al.
Published: (2024)
by: Ding, Tianyu, et al.
Published: (2024)
Lisbon Computational Linguists at SemEval-2024 Task 2: Using A Mistral 7B Model and Data Augmentation
by: Guimarães, Artur, et al.
Published: (2024)
by: Guimarães, Artur, et al.
Published: (2024)
Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
by: Dosajh, Arjun, et al.
Published: (2025)
by: Dosajh, Arjun, et al.
Published: (2025)
Mixture of Experts Approaches in Dense Retrieval Tasks
by: Sokli, Effrosyni, et al.
Published: (2025)
by: Sokli, Effrosyni, et al.
Published: (2025)
Prompt Perturbation in Retrieval-Augmented Generation based Large Language Models
by: Hu, Zhibo, et al.
Published: (2024)
by: Hu, Zhibo, et al.
Published: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
by: Preuveneers, Jack, et al.
Published: (2025)
by: Preuveneers, Jack, et al.
Published: (2025)
BLP-2023 Task 2: Sentiment Analysis
by: Hasan, Md. Arid, et al.
Published: (2023)
by: Hasan, Md. Arid, et al.
Published: (2023)
MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
by: Wei, Kaiwen, et al.
Published: (2025)
by: Wei, Kaiwen, et al.
Published: (2025)
Efficacy of a Computer Tutor that Models Expert Human Tutors
by: Olney, Andrew M., et al.
Published: (2025)
by: Olney, Andrew M., et al.
Published: (2025)
TrustAI at SemEval-2024 Task 8: A Comprehensive Analysis of Multi-domain Machine Generated Text Detection Techniques
by: Urlana, Ashok, et al.
Published: (2024)
by: Urlana, Ashok, et al.
Published: (2024)
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
by: Johnson, Warren, et al.
Published: (2026)
by: Johnson, Warren, et al.
Published: (2026)
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
by: Bhaskar, Yash, et al.
Published: (2025)
by: Bhaskar, Yash, et al.
Published: (2025)
HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction
by: Arief, Hasan
Published: (2026)
by: Arief, Hasan
Published: (2026)
Similar Items
-
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
by: Ahmad, Sarfraz, et al.
Published: (2025) -
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025) -
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
by: Bouthors, Maxime, et al.
Published: (2025) -
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
by: Nzeyimana, Antoine, et al.
Published: (2025) -
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)