Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
Fuente:
arXiv
Saved in:
| Main Authors: | Gema, Aryo Pradipta, Lee, Chaeeun, Minervini, Pasquale, Daines, Luke, Simpson, T. Ian, Alex, Beatrice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023)
by: Gema, Aryo Pradipta, et al.
Published: (2023)
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Self-Training Large Language Models for Tool-Use Without Demonstrations
by: Luo, Ne, et al.
Published: (2025)
by: Luo, Ne, et al.
Published: (2025)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
by: Kwan, Wai-Chung, et al.
Published: (2026)
by: Kwan, Wai-Chung, et al.
Published: (2026)
IryoNLP at MEDIQA-CORR 2024: Tackling the Medical Error Detection & Correction Task On the Shoulders of Medical Agents
by: Corbeil, Jean-Philippe
Published: (2024)
by: Corbeil, Jean-Philippe
Published: (2024)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Can GPT-3.5 Generate and Code Discharge Summaries?
by: Falis, Matúš, et al.
Published: (2024)
by: Falis, Matúš, et al.
Published: (2024)
Process-Supervised Multi-Agent Reinforcement Learning for Reliable Clinical Reasoning
by: Lee, Chaeeun, et al.
Published: (2026)
by: Lee, Chaeeun, et al.
Published: (2026)
MediFact at MEDIQA-CORR 2024: Why AI Needs a Human Touch
by: Saeed, Nadia
Published: (2024)
by: Saeed, Nadia
Published: (2024)
PromptMind Team at MEDIQA-CORR 2024: Improving Clinical Text Correction with Error Categorization and LLM Ensembles
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
WangLab at MEDIQA-CORR 2024: Optimized LLM-based Programs for Medical Error Detection and Correction
by: Toma, Augustin, et al.
Published: (2024)
by: Toma, Augustin, et al.
Published: (2024)
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
by: Hong, Giwon, et al.
Published: (2024)
by: Hong, Giwon, et al.
Published: (2024)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
by: Murphy, Alexander, et al.
Published: (2025)
by: Murphy, Alexander, et al.
Published: (2025)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
A Comparative Study on Patient Language across Therapeutic Domains for Effective Patient Voice Classification in Online Health Discussions
by: Lysandrou, Giorgos, et al.
Published: (2024)
by: Lysandrou, Giorgos, et al.
Published: (2024)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
by: Rajani, Neel, et al.
Published: (2025)
by: Rajani, Neel, et al.
Published: (2025)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025)
by: Gema, Aryo Pradipta, et al.
Published: (2025)
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
by: Leang, Joshua Ong Jun, et al.
Published: (2025)
by: Leang, Joshua Ong Jun, et al.
Published: (2025)
GRADA: Graph-based Reranking against Adversarial Documents Attack
by: Zheng, Jingjie, et al.
Published: (2025)
by: Zheng, Jingjie, et al.
Published: (2025)
WangLab at MEDIQA-M3G 2024: Multimodal Medical Answer Generation using Large Language Models
by: Xie, Ronald, et al.
Published: (2024)
by: Xie, Ronald, et al.
Published: (2024)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
by: Attimonelli, Matteo, et al.
Published: (2026)
by: Attimonelli, Matteo, et al.
Published: (2026)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
UMass-BioNLP at MEDIQA-M3G 2024: DermPrompt -- A Systematic Exploration of Prompt Engineering with GPT-4V for Dermatological Diagnosis
by: Vashisht, Parth, et al.
Published: (2024)
by: Vashisht, Parth, et al.
Published: (2024)
FairBelief -- Assessing Harmful Beliefs in Language Models
by: Setzu, Mattia, et al.
Published: (2024)
by: Setzu, Mattia, et al.
Published: (2024)
D-NLP at SemEval-2024 Task 2: Evaluating Clinical Inference Capabilities of Large Language Models
by: Altinok, Duygu
Published: (2024)
by: Altinok, Duygu
Published: (2024)
Answerability in Retrieval-Augmented Open-Domain Question Answering
by: Abdumalikov, Rustam, et al.
Published: (2024)
by: Abdumalikov, Rustam, et al.
Published: (2024)
Evaluating and Adapting Large Language Models to Represent Folktales in Low-Resource Languages
by: Meaney, JA, et al.
Published: (2024)
by: Meaney, JA, et al.
Published: (2024)
How Well Do Large Language Models Truly Ground?
by: Lee, Hyunji, et al.
Published: (2023)
by: Lee, Hyunji, et al.
Published: (2023)
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
by: Ghazaryan, Gayane, et al.
Published: (2024)
by: Ghazaryan, Gayane, et al.
Published: (2024)
Progressive-Hint Prompting Improves Reasoning in Large Language Models
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
SPARSEFIT: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations
by: Solano, Jesus, et al.
Published: (2023)
by: Solano, Jesus, et al.
Published: (2023)
Active Learning for NLP with Large Language Models
by: Wang, Xuesong
Published: (2024)
by: Wang, Xuesong
Published: (2024)
Using Natural Language Explanations to Improve Robustness of In-context Learning
by: He, Xuanli, et al.
Published: (2023)
by: He, Xuanli, et al.
Published: (2023)
NavHint: Vision and Language Navigation Agent with a Hint Generator
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Are We Done with MMLU?
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Similar Items
-
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
by: Gema, Aryo Pradipta, et al.
Published: (2024) -
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023) -
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025) -
Self-Training Large Language Models for Tool-Use Without Demonstrations
by: Luo, Ne, et al.
Published: (2025) -
Noiser: Bounded Input Perturbations for Attributing Large Language Models
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)