Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset
Fuente:
arXiv
Salvato in:
| Autori principali: | Rasool, Zafaryab, Kurniawan, Stefanus, Balugo, Sherwin, Barnett, Scott, Vasa, Rajesh, Chesser, Courtney, Hampstead, Benjamin M., Belleville, Sylvie, Mouzakis, Kon, Bahar-Fuchs, Alex |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RAGProbe: An Automated Approach for Evaluating RAG Applications
di: Sivasothy, Shangeetha, et al.
Pubblicazione: (2024)
di: Sivasothy, Shangeetha, et al.
Pubblicazione: (2024)
LLMs for Test Input Generation for Semantic Caches
di: Rasool, Zafaryab, et al.
Pubblicazione: (2024)
di: Rasool, Zafaryab, et al.
Pubblicazione: (2024)
The M-factor: A Novel Metric for Evaluating Neural Architecture Search in Resource-Constrained Environments
di: Thudumu, Srikanth, et al.
Pubblicazione: (2025)
di: Thudumu, Srikanth, et al.
Pubblicazione: (2025)
A Survey on Context-Aware Multi-Agent Systems: Techniques, Challenges and Future Directions
di: Du, Hung, et al.
Pubblicazione: (2024)
di: Du, Hung, et al.
Pubblicazione: (2024)
Contextual Knowledge Sharing in Multi-Agent Reinforcement Learning with Decentralized Communication and Coordination
di: Du, Hung, et al.
Pubblicazione: (2025)
di: Du, Hung, et al.
Pubblicazione: (2025)
Goal-Oriented Multi-Agent Reinforcement Learning for Decentralized Agent Teams
di: Du, Hung, et al.
Pubblicazione: (2025)
di: Du, Hung, et al.
Pubblicazione: (2025)
Overcoming Semantic Dilution in Transformer-Based Next Frame Prediction
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
CSAOT: Cooperative Multi-Agent System for Active Object Tracking
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
Local Control Networks (LCNs): Optimizing Flexibility in Neural Network Data Pattern Capture
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
Fine-Tuning or Fine-Failing? Debunking Performance Myths in Large Language Models
di: Barnett, Scott, et al.
Pubblicazione: (2024)
di: Barnett, Scott, et al.
Pubblicazione: (2024)
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
di: Widanapathiranage, Dilani, et al.
Pubblicazione: (2025)
di: Widanapathiranage, Dilani, et al.
Pubblicazione: (2025)
Seven Failure Points When Engineering a Retrieval Augmented Generation System
di: Barnett, Scott, et al.
Pubblicazione: (2024)
di: Barnett, Scott, et al.
Pubblicazione: (2024)
Dual-Branch HNSW Approach with Skip Bridges and LID-Driven Optimization
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
di: Nguyen, Hy, et al.
Pubblicazione: (2025)
Large language models for generating rules, yay or nay?
di: Sivasothy, Shangeetha, et al.
Pubblicazione: (2024)
di: Sivasothy, Shangeetha, et al.
Pubblicazione: (2024)
ML-On-Rails: Safeguarding Machine Learning Models in Software Systems A Case Study
di: Abdelkader, Hala, et al.
Pubblicazione: (2024)
di: Abdelkader, Hala, et al.
Pubblicazione: (2024)
When loxodromics are pseudo-Anosovs on witnesses
di: Chesser, Marissa
Pubblicazione: (2026)
di: Chesser, Marissa
Pubblicazione: (2026)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
di: Pattnayak, Priyaranjan, et al.
Pubblicazione: (2025)
WOOD MACHINING PROPERTIES OF AUSTRALIAN PLANTATION-GROWN EUCALYPTS
di: Benoit Belleville
Pubblicazione: (2016)
di: Benoit Belleville
Pubblicazione: (2016)
WOOD PLANING PROPERTIES OF AUSTRALIAN PLANTATION-GROWN Eucalypts
di: Benoit Belleville
Pubblicazione: (2016)
di: Benoit Belleville
Pubblicazione: (2016)
Gluing characteristics of Papua New Guinea timber species for various non-structural applications
di: Benoit Belleville
Pubblicazione: (2024)
di: Benoit Belleville
Pubblicazione: (2024)
Assessment of physical and mechanical properties of Papua New Guinea timber species
di: Benoit Belleville
Pubblicazione: (2020)
di: Benoit Belleville
Pubblicazione: (2020)
Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
di: Basem, Mohamed, et al.
Pubblicazione: (2025)
di: Basem, Mohamed, et al.
Pubblicazione: (2025)
Numerical Analysis of Lensless Imaging with Active Metasurfaces and Single-Pixel Detectors
di: Belleville, Julie, et al.
Pubblicazione: (2024)
di: Belleville, Julie, et al.
Pubblicazione: (2024)
Scientific QA System with Verifiable Answers
di: Ljajić, Adela, et al.
Pubblicazione: (2024)
di: Ljajić, Adela, et al.
Pubblicazione: (2024)
Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
di: Mozafari, Jamshid, et al.
Pubblicazione: (2025)
di: Mozafari, Jamshid, et al.
Pubblicazione: (2025)
NLP at UC Santa Cruz at SemEval-2024 Task 5: Legal Answer Validation using Few-Shot Multi-Choice QA
di: Pahilajani, Anish, et al.
Pubblicazione: (2024)
di: Pahilajani, Anish, et al.
Pubblicazione: (2024)
Explainablity QA dataset
di: Anonymous
Pubblicazione: (2026)
di: Anonymous
Pubblicazione: (2026)
Educational attainment mitigates hippocampal‐related episodic memory decline in individuals at risk of Alzheimer's disease
di: Annalise Aleta LaPlume, et al.
Pubblicazione: (2026)
di: Annalise Aleta LaPlume, et al.
Pubblicazione: (2026)
RealTime QA: What's the Answer Right Now?
di: Kasai, Jungo, et al.
Pubblicazione: (2022)
di: Kasai, Jungo, et al.
Pubblicazione: (2022)
ExpertQA: Expert-Curated Questions and Attributed Answers
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
di: Malaviya, Chaitanya, et al.
Pubblicazione: (2023)
Volume Tracking Based Reference Mesh Extraction for Time-Varying Mesh Compression
di: Chen, Guodong, et al.
Pubblicazione: (2024)
di: Chen, Guodong, et al.
Pubblicazione: (2024)
Purely pseudo-Anosov subgroups of the genus two handlebody group
di: Chesser, Marissa, et al.
Pubblicazione: (2023)
di: Chesser, Marissa, et al.
Pubblicazione: (2023)
EEE-QA: Exploring Effective and Efficient Question-Answer Representations
di: Hu, Zhanghao, et al.
Pubblicazione: (2024)
di: Hu, Zhanghao, et al.
Pubblicazione: (2024)
Optimal Differentially Private Sampling of Unbounded Gaussians
di: Iverson, Valentio, et al.
Pubblicazione: (2025)
di: Iverson, Valentio, et al.
Pubblicazione: (2025)
Return of EM: Entity-driven Answer Set Expansion for QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
MentalQA: An Annotated Arabic Corpus for Questions and Answers of Mental Healthcare
di: Alhuzali, Hassan, et al.
Pubblicazione: (2024)
di: Alhuzali, Hassan, et al.
Pubblicazione: (2024)
Ensuring Robustness in ML-enabled Software Systems: A User Survey
di: Abdelkader, Hala, et al.
Pubblicazione: (2025)
di: Abdelkader, Hala, et al.
Pubblicazione: (2025)
A taxonomy of grain boundary migration mechanisms via displacement texture characterization
di: Chesser, Ian, et al.
Pubblicazione: (2021)
di: Chesser, Ian, et al.
Pubblicazione: (2021)
Decentralized Multi-product Pricing: Diagonal Dominance, Nash Equilibrium, and Price of Anarchy
di: Chen, Boxiao, et al.
Pubblicazione: (2026)
di: Chen, Boxiao, et al.
Pubblicazione: (2026)
SWE-QA: Can Language Models Answer Repository-level Code Questions?
di: Peng, Weihan, et al.
Pubblicazione: (2025)
di: Peng, Weihan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RAGProbe: An Automated Approach for Evaluating RAG Applications
di: Sivasothy, Shangeetha, et al.
Pubblicazione: (2024) -
LLMs for Test Input Generation for Semantic Caches
di: Rasool, Zafaryab, et al.
Pubblicazione: (2024) -
The M-factor: A Novel Metric for Evaluating Neural Architecture Search in Resource-Constrained Environments
di: Thudumu, Srikanth, et al.
Pubblicazione: (2025) -
A Survey on Context-Aware Multi-Agent Systems: Techniques, Challenges and Future Directions
di: Du, Hung, et al.
Pubblicazione: (2024) -
Contextual Knowledge Sharing in Multi-Agent Reinforcement Learning with Decentralized Communication and Coordination
di: Du, Hung, et al.
Pubblicazione: (2025)