Saved in:
| Main Authors: | Capdehourat, Germán, Amigo, Isabel, Lorenzo, Brian, Trigo, Joaquín |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.18072 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Writing as a testbed for open ended agents
by: Gooding, Sian, et al.
Published: (2025)
by: Gooding, Sian, et al.
Published: (2025)
Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question Answering
by: Wen, Zhihua, et al.
Published: (2024)
by: Wen, Zhihua, et al.
Published: (2024)
ClinText-SP and RigoBERTa Clinical: a new set of open resources for Spanish Clinical NLP
by: Subies, Guillem García, et al.
Published: (2025)
by: Subies, Guillem García, et al.
Published: (2025)
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
by: Luo, Yifan, et al.
Published: (2024)
by: Luo, Yifan, et al.
Published: (2024)
RealMedQA: A pilot biomedical question answering dataset containing realistic clinical questions
by: Kell, Gregory, et al.
Published: (2024)
by: Kell, Gregory, et al.
Published: (2024)
Generative AI for automatic topic labelling
by: Kozlowski, Diego, et al.
Published: (2024)
by: Kozlowski, Diego, et al.
Published: (2024)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
by: Maar, Jim, et al.
Published: (2026)
by: Maar, Jim, et al.
Published: (2026)
FMI@SU ToxHabits: Evaluating LLMs Performance on Toxic Habit Extraction in Spanish Clinical Texts
by: Vassileva, Sylvia, et al.
Published: (2026)
by: Vassileva, Sylvia, et al.
Published: (2026)
Why does in-context learning fail sometimes? Evaluating in-context learning on open and closed questions
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
Can formal argumentative reasoning enhance LLMs performances?
by: Castagna, Federico, et al.
Published: (2024)
by: Castagna, Federico, et al.
Published: (2024)
The 20 questions game to distinguish large language models
by: Richardeau, Gurvan, et al.
Published: (2024)
by: Richardeau, Gurvan, et al.
Published: (2024)
A cross-species neural foundation model for end-to-end speech decoding
by: Zhang, Yizi, et al.
Published: (2025)
by: Zhang, Yizi, et al.
Published: (2025)
Seventeenth-Century Spanish American Notary Records for Fine-Tuning Spanish Large Language Models
by: Sarker, Shraboni, et al.
Published: (2024)
by: Sarker, Shraboni, et al.
Published: (2024)
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
by: Maharjan, Jenish, et al.
Published: (2024)
by: Maharjan, Jenish, et al.
Published: (2024)
$How^{2}$: How to learn from procedural How-to questions
by: Dagan, Gautier, et al.
Published: (2025)
by: Dagan, Gautier, et al.
Published: (2025)
Using LLMs to identify features of personal and professional skills in an open-response situational judgment test
by: Walsh, Cole, et al.
Published: (2025)
by: Walsh, Cole, et al.
Published: (2025)
Research on emotionally intelligent dialogue generation based on automatic dialogue system
by: Wang, Jin, et al.
Published: (2024)
by: Wang, Jin, et al.
Published: (2024)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
by: Plaza, Irene, et al.
Published: (2024)
by: Plaza, Irene, et al.
Published: (2024)
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
by: Buscemi, Alessio, et al.
Published: (2025)
by: Buscemi, Alessio, et al.
Published: (2025)
AutoHarness: improving LLM agents by automatically synthesizing a code harness
by: Lou, Xinghua, et al.
Published: (2026)
by: Lou, Xinghua, et al.
Published: (2026)
Evaluating Students' Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large
by: Jauhiainen, Jussi S., et al.
Published: (2024)
by: Jauhiainen, Jussi S., et al.
Published: (2024)
Building another Spanish dictionary, this time with GPT-4
by: Ortega-Martín, Miguel, et al.
Published: (2024)
by: Ortega-Martín, Miguel, et al.
Published: (2024)
Agribot: agriculture-specific question answer system
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
keqing: knowledge-based question answering is a nature chain-of-thought mentor of LLM
by: Wang, Chaojie, et al.
Published: (2023)
by: Wang, Chaojie, et al.
Published: (2023)
From text to multimodal: a survey of adversarial example generation in question answering systems
by: Yigit, Gulsum, et al.
Published: (2023)
by: Yigit, Gulsum, et al.
Published: (2023)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
Developing an AI framework to automatically detect shared decision-making in patient-doctor conversations
by: Ponce-Ponte, Oscar J., et al.
Published: (2025)
by: Ponce-Ponte, Oscar J., et al.
Published: (2025)
Automated evaluation of LLMs for effective machine translation of Mandarin Chinese to English
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Online Social Support Detection in Spanish Social Media Texts
by: Tash, Moein Shahiki, et al.
Published: (2025)
by: Tash, Moein Shahiki, et al.
Published: (2025)
Spanish TrOCR: Leveraging Transfer Learning for Language Adaptation
by: Lauar, Filipe, et al.
Published: (2024)
by: Lauar, Filipe, et al.
Published: (2024)
NoticIA: A Clickbait Article Summarization Dataset in Spanish
by: García-Ferrero, Iker, et al.
Published: (2024)
by: García-Ferrero, Iker, et al.
Published: (2024)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
by: Cunha, Washington, et al.
Published: (2025)
by: Cunha, Washington, et al.
Published: (2025)
Comparative analysis of privacy-preserving open-source LLMs regarding extraction of diagnostic information from clinical CMR imaging reports
by: Amirrajab, Sina, et al.
Published: (2025)
by: Amirrajab, Sina, et al.
Published: (2025)
ACL-Verbatim: hallucination-free question answering for research
by: Recski, Gábor, et al.
Published: (2026)
by: Recski, Gábor, et al.
Published: (2026)
Founder effects shape the evolutionary dynamics of multimodality in open LLM families
by: Cebrian, Manuel
Published: (2026)
by: Cebrian, Manuel
Published: (2026)
MessIRve: A Large-Scale Spanish Information Retrieval Dataset
by: Valentini, Francisco, et al.
Published: (2024)
by: Valentini, Francisco, et al.
Published: (2024)
Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
by: Schaaff, Kristina, et al.
Published: (2023)
by: Schaaff, Kristina, et al.
Published: (2023)
SciDER: Scientific Data-centric End-to-end Researcher
by: Lin, Ke, et al.
Published: (2026)
by: Lin, Ke, et al.
Published: (2026)
Similar Items
-
Writing as a testbed for open ended agents
by: Gooding, Sian, et al.
Published: (2025) -
Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question Answering
by: Wen, Zhihua, et al.
Published: (2024) -
ClinText-SP and RigoBERTa Clinical: a new set of open resources for Spanish Clinical NLP
by: Subies, Guillem García, et al.
Published: (2025) -
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
by: Luo, Yifan, et al.
Published: (2024) -
RealMedQA: A pilot biomedical question answering dataset containing realistic clinical questions
by: Kell, Gregory, et al.
Published: (2024)