Benchmarking LLMs and SLMs for patient reported outcomes
Fuente:
arXiv
Saved in:
| Main Authors: | Marengo, Matteo, Lévy, Jarod, Bibault, Jean-Emmanuel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
G-Boost: Boosting Private SLMs with General LLMs
by: Fan, Yijiang, et al.
Published: (2025)
by: Fan, Yijiang, et al.
Published: (2025)
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
by: Marco, Guillermo, et al.
Published: (2024)
by: Marco, Guillermo, et al.
Published: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
by: Bhola, Ishaan, et al.
Published: (2025)
by: Bhola, Ishaan, et al.
Published: (2025)
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
Return of the Encoder: Maximizing Parameter Efficiency for SLMs
by: Elfeki, Mohamed, et al.
Published: (2025)
by: Elfeki, Mohamed, et al.
Published: (2025)
DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs
by: Hasan, Md Hasebul, et al.
Published: (2026)
by: Hasan, Md Hasebul, et al.
Published: (2026)
CCQA: Generating Question from Solution Can Improve Inference-Time Reasoning in SLMs
by: Kim, Jin Young, et al.
Published: (2025)
by: Kim, Jin Young, et al.
Published: (2025)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
by: Chen, Jennifer, et al.
Published: (2025)
by: Chen, Jennifer, et al.
Published: (2025)
A Structure-Agnostic Co-Tuning Framework for LLMs and SLMs in Cloud-Edge Systems
by: Liu, Yuze, et al.
Published: (2025)
by: Liu, Yuze, et al.
Published: (2025)
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
by: Padarha, Shreyansh
Published: (2025)
by: Padarha, Shreyansh
Published: (2025)
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs
by: Alrashed, Sultan
Published: (2024)
by: Alrashed, Sultan
Published: (2024)
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
by: Liang, Zhuowen, et al.
Published: (2026)
by: Liang, Zhuowen, et al.
Published: (2026)
EasyMath: A 0-shot Math Benchmark for SLMs
by: Karki, Drishya, et al.
Published: (2025)
by: Karki, Drishya, et al.
Published: (2025)
CEGI: Measuring the trade-off between efficiency and carbon emissions for SLMs and VLMs
by: Kumar, Abhas, et al.
Published: (2024)
by: Kumar, Abhas, et al.
Published: (2024)
Brain-to-Text Decoding: A Non-invasive Approach via Typing
by: Lévy, Jarod, et al.
Published: (2025)
by: Lévy, Jarod, et al.
Published: (2025)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
by: Tie, Guiyao, et al.
Published: (2025)
by: Tie, Guiyao, et al.
Published: (2025)
Fine-Tuning Small Language Models (SLMs) for Autonomous Web-based Geographical Information Systems (AWebGIS)
by: Ashani, Mahdi Nazari, et al.
Published: (2025)
by: Ashani, Mahdi Nazari, et al.
Published: (2025)
Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
by: Boffa, Matteo, et al.
Published: (2025)
by: Boffa, Matteo, et al.
Published: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
DHP Benchmark: Are LLMs Good NLG Evaluators?
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
Flames: Benchmarking Value Alignment of LLMs in Chinese
by: Huang, Kexin, et al.
Published: (2023)
by: Huang, Kexin, et al.
Published: (2023)
WebWalker: Benchmarking LLMs in Web Traversal
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
by: Huang, Xu, et al.
Published: (2025)
by: Huang, Xu, et al.
Published: (2025)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
Benchmarking LLMs for Mimicking Child-Caregiver Language in Interaction
by: Liu, Jing, et al.
Published: (2024)
by: Liu, Jing, et al.
Published: (2024)
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
by: Zheng, Huaixiu Steven, et al.
Published: (2024)
by: Zheng, Huaixiu Steven, et al.
Published: (2024)
Benchmarking and Adapting On-Device LLMs for Clinical Decision Support
by: Munim, Alif, et al.
Published: (2025)
by: Munim, Alif, et al.
Published: (2025)
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
by: Ajjour, Yamen, et al.
Published: (2026)
by: Ajjour, Yamen, et al.
Published: (2026)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
Beyond Specialization: Benchmarking LLMs for Transliteration of Indian Languages
by: Azam, Gulfarogh, et al.
Published: (2025)
by: Azam, Gulfarogh, et al.
Published: (2025)
When Benchmarks Leak: Inference-Time Decontamination for LLMs
by: Chai, Jianzhe, et al.
Published: (2026)
by: Chai, Jianzhe, et al.
Published: (2026)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Evaluating LLMs on Entity Disambiguation in Tables
by: Belotti, Federico, et al.
Published: (2024)
by: Belotti, Federico, et al.
Published: (2024)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
by: Zhou, Yitong, et al.
Published: (2025)
by: Zhou, Yitong, et al.
Published: (2025)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
by: Hong, Zijin, et al.
Published: (2025)
by: Hong, Zijin, et al.
Published: (2025)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
by: Lei, Chao, et al.
Published: (2025)
by: Lei, Chao, et al.
Published: (2025)
Similar Items
-
G-Boost: Boosting Private SLMs with General LLMs
by: Fan, Yijiang, et al.
Published: (2025) -
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs
by: Marco, Guillermo, et al.
Published: (2024) -
Orca-Math: Unlocking the potential of SLMs in Grade School Math
by: Mitra, Arindam, et al.
Published: (2024) -
Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
by: Bhola, Ishaan, et al.
Published: (2025) -
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
by: Corbeil, Jean-Philippe, et al.
Published: (2025)