Guardado en:
| Autores principales: | Bai, Fan, Harrigian, Keith, Stremmel, Joel, Hassanzadeh, Hamid, Saeedi, Ardavan, Dredze, Mark |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2412.04573 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition
por: Bai, Fan, et al.
Publicado: (2025)
por: Bai, Fan, et al.
Publicado: (2025)
Are Clinical T5 Models Better for Clinical Text?
por: Li, Yahan, et al.
Publicado: (2024)
por: Li, Yahan, et al.
Publicado: (2024)
Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2026)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2026)
Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict
por: Sun, Kaiser, et al.
Publicado: (2025)
por: Sun, Kaiser, et al.
Publicado: (2025)
Consistency Training by Synthetic Question Generation for Conversational Question Answering
por: Hemati, Hamed Hematian, et al.
Publicado: (2024)
por: Hemati, Hamed Hematian, et al.
Publicado: (2024)
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
por: Chen, Hanjie, et al.
Publicado: (2024)
por: Chen, Hanjie, et al.
Publicado: (2024)
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
por: An, Bang, et al.
Publicado: (2025)
por: An, Bang, et al.
Publicado: (2025)
Amuro and Char: Analyzing the Relationship between Pre-Training and Fine-Tuning of Large Language Models
por: Sun, Kaiser, et al.
Publicado: (2024)
por: Sun, Kaiser, et al.
Publicado: (2024)
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
por: Wanner, Miriam, et al.
Publicado: (2024)
por: Wanner, Miriam, et al.
Publicado: (2024)
Schema-Driven Information Extraction from Heterogeneous Tables
por: Bai, Fan, et al.
Publicado: (2023)
por: Bai, Fan, et al.
Publicado: (2023)
Evaluating Biases in Context-Dependent Health Questions
por: Levy, Sharon, et al.
Publicado: (2024)
por: Levy, Sharon, et al.
Publicado: (2024)
Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets
por: Daswani, Ashwin, et al.
Publicado: (2024)
por: Daswani, Ashwin, et al.
Publicado: (2024)
Can one size fit all?: Measuring Failure in Multi-Document Summarization Domain Transfer
por: DeLucia, Alexandra, et al.
Publicado: (2025)
por: DeLucia, Alexandra, et al.
Publicado: (2025)
Evaluating the Evaluators: Are readability metrics good measures of readability?
por: Cachola, Isabel, et al.
Publicado: (2025)
por: Cachola, Isabel, et al.
Publicado: (2025)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
por: Jahara, Fatima, et al.
Publicado: (2025)
por: Jahara, Fatima, et al.
Publicado: (2025)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
por: Sun, Kaiser, et al.
Publicado: (2026)
por: Sun, Kaiser, et al.
Publicado: (2026)
Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language Model
por: Kim, Taehee, et al.
Publicado: (2024)
por: Kim, Taehee, et al.
Publicado: (2024)
LLMs in Biomedicine: A study on clinical Named Entity Recognition
por: Monajatipoor, Masoud, et al.
Publicado: (2024)
por: Monajatipoor, Masoud, et al.
Publicado: (2024)
From Policy to Logic for Efficient and Interpretable Coverage Assessment
por: Pokharel, Rhitabrat, et al.
Publicado: (2026)
por: Pokharel, Rhitabrat, et al.
Publicado: (2026)
ExpertQA: Expert-Curated Questions and Attributed Answers
por: Malaviya, Chaitanya, et al.
Publicado: (2023)
por: Malaviya, Chaitanya, et al.
Publicado: (2023)
Towards Better Question Generation in QA-based Event Extraction
por: Hong, Zijin, et al.
Publicado: (2024)
por: Hong, Zijin, et al.
Publicado: (2024)
Weird Generalization is Weirdly Brittle
por: Wanner, Miriam, et al.
Publicado: (2026)
por: Wanner, Miriam, et al.
Publicado: (2026)
NeoQA: Evidence-based Question Answering with Generated News Events
por: Glockner, Max, et al.
Publicado: (2025)
por: Glockner, Max, et al.
Publicado: (2025)
Prompting-based Synthetic Data Generation for Few-Shot Question Answering
por: Schmidt, Maximilian, et al.
Publicado: (2024)
por: Schmidt, Maximilian, et al.
Publicado: (2024)
SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation
por: Kundu, Debarshi
Publicado: (2024)
por: Kundu, Debarshi
Publicado: (2024)
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
por: Yifei, Li S., et al.
Publicado: (2025)
por: Yifei, Li S., et al.
Publicado: (2025)
PolQA: Polish Question Answering Dataset
por: Rybak, Piotr, et al.
Publicado: (2022)
por: Rybak, Piotr, et al.
Publicado: (2022)
Building Open-Retrieval Conversational Question Answering Systems by Generating Synthetic Data and Decontextualizing User Questions
por: Vlachos, Christos, et al.
Publicado: (2025)
por: Vlachos, Christos, et al.
Publicado: (2025)
Synthetic Context Generation for Question Generation
por: Liu, Naiming, et al.
Publicado: (2024)
por: Liu, Naiming, et al.
Publicado: (2024)
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats
por: Sasse, Kuleen, et al.
Publicado: (2024)
por: Sasse, Kuleen, et al.
Publicado: (2024)
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
por: Huang, Heyuan, et al.
Publicado: (2025)
por: Huang, Heyuan, et al.
Publicado: (2025)
JDocQA: Japanese Document Question Answering Dataset for Generative Language Models
por: Onami, Eri, et al.
Publicado: (2024)
por: Onami, Eri, et al.
Publicado: (2024)
On the Failure of Latent State Persistence in Large Language Models
por: Huang, Jen-tse, et al.
Publicado: (2025)
por: Huang, Jen-tse, et al.
Publicado: (2025)
Assessing The Potential Of Mid-Sized Language Models For Clinical QA
por: Bolton, Elliot, et al.
Publicado: (2024)
por: Bolton, Elliot, et al.
Publicado: (2024)
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
por: Schimanski, Tobias, et al.
Publicado: (2026)
por: Schimanski, Tobias, et al.
Publicado: (2026)
DebateQA: Evaluating Question Answering on Debatable Knowledge
por: Xu, Rongwu, et al.
Publicado: (2024)
por: Xu, Rongwu, et al.
Publicado: (2024)
A Closer Look at Claim Decomposition
por: Wanner, Miriam, et al.
Publicado: (2024)
por: Wanner, Miriam, et al.
Publicado: (2024)
Give me a hint: Can LLMs take a hint to solve math problems?
por: Agrawal, Vansh, et al.
Publicado: (2024)
por: Agrawal, Vansh, et al.
Publicado: (2024)
Synthetic Multimodal Question Generation
por: Wu, Ian, et al.
Publicado: (2024)
por: Wu, Ian, et al.
Publicado: (2024)
Improving Clinical NLP Performance through Language Model-Generated Synthetic Clinical Data
por: Chen, Shan, et al.
Publicado: (2024)
por: Chen, Shan, et al.
Publicado: (2024)
Ejemplares similares
-
LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition
por: Bai, Fan, et al.
Publicado: (2025) -
Are Clinical T5 Models Better for Clinical Text?
por: Li, Yahan, et al.
Publicado: (2024) -
Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2026) -
Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict
por: Sun, Kaiser, et al.
Publicado: (2025) -
Consistency Training by Synthetic Question Generation for Conversational Question Answering
por: Hemati, Hamed Hematian, et al.
Publicado: (2024)