CaLMQA: Exploring culturally specific long-form question answering across 23 languages
Fuente:
arXiv
Saved in:
| Main Authors: | Arora, Shane, Karpinska, Marzena, Chen, Hung-Ting, Bhattacharjee, Ipsita, Iyyer, Mohit, Choi, Eunsol |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025)
by: Kim, Yekyung, et al.
Published: (2025)
One Thousand and One Pairs: A "novel" challenge for long-context language models
by: Karpinska, Marzena, et al.
Published: (2024)
by: Karpinska, Marzena, et al.
Published: (2024)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025)
by: Mekala, Anmol, et al.
Published: (2025)
Understanding Retrieval Augmentation for Long-Form Question Answering
by: Chen, Hung-Ting, et al.
Published: (2023)
by: Chen, Hung-Ting, et al.
Published: (2023)
OverThink: Slowdown Attacks on Reasoning LLMs
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
AI use in American newspapers is widespread, uneven, and rarely disclosed
by: Russell, Jenna, et al.
Published: (2025)
by: Russell, Jenna, et al.
Published: (2025)
VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
by: Song, Yixiao, et al.
Published: (2024)
by: Song, Yixiao, et al.
Published: (2024)
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature
by: Srivastava, Alisha, et al.
Published: (2025)
by: Srivastava, Alisha, et al.
Published: (2025)
FABLES: Evaluating faithfulness and content selection in book-length summarization
by: Kim, Yekyung, et al.
Published: (2024)
by: Kim, Yekyung, et al.
Published: (2024)
Frankentext: Stitching random text fragments into long-form narratives
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
Open-World Evaluation for Retrieving Diverse Perspectives
by: Chen, Hung-Ting, et al.
Published: (2024)
by: Chen, Hung-Ting, et al.
Published: (2024)
VeriFastScore: Speeding up long-form factuality evaluation
by: Rajendhran, Rishanth, et al.
Published: (2025)
by: Rajendhran, Rishanth, et al.
Published: (2025)
Agribot: agriculture-specific question answer system
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
Suri: Multi-constraint Instruction Following for Long-form Text Generation
by: Pham, Chau Minh, et al.
Published: (2024)
by: Pham, Chau Minh, et al.
Published: (2024)
RVR: Retrieve-Verify-Retrieve for Comprehensive Question Answering
by: Qian, Deniz, et al.
Published: (2026)
by: Qian, Deniz, et al.
Published: (2026)
Localizing and Mitigating Errors in Long-form Question Answering
by: Sachdeva, Rachneet, et al.
Published: (2024)
by: Sachdeva, Rachneet, et al.
Published: (2024)
Literary Evidence Retrieval via Long-Context Language Models
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
Mailbag questions and answers
by: Richard Rainsberger
Published: (2025)
by: Richard Rainsberger
Published: (2025)
100 questions answered?
Mailbag questions and answers
by: Richard Rainsberger
Published: (2025)
by: Richard Rainsberger
Published: (2025)
Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
by: Chen, Hung-Ting, et al.
Published: (2025)
by: Chen, Hung-Ting, et al.
Published: (2025)
70B-parameter large language models in Japanese medical question-answering
by: Sukeda, Issey, et al.
Published: (2024)
by: Sukeda, Issey, et al.
Published: (2024)
Large language models provide unsafe answers to patient-posed medical questions
by: Draelos, Rachel L., et al.
Published: (2025)
by: Draelos, Rachel L., et al.
Published: (2025)
Performance of large language models in answering frequently‐asked questions on celiac disease
by: Nadav Peled, et al.
Published: (2026)
by: Nadav Peled, et al.
Published: (2026)
RefreshKV: Updating Small KV Cache During Long-form Generation
by: Xu, Fangyuan, et al.
Published: (2024)
by: Xu, Fangyuan, et al.
Published: (2024)
An answer to the question for the analytical device
by: Jairo Báez
Published: (2010)
by: Jairo Báez
Published: (2010)
AmbigDocs: Reasoning across Documents on Different Entities under the Same Name
by: Lee, Yoonsang, et al.
Published: (2024)
by: Lee, Yoonsang, et al.
Published: (2024)
Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation
by: Jafari, Nazanin, et al.
Published: (2026)
by: Jafari, Nazanin, et al.
Published: (2026)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
by: Samuel, Vinay, et al.
Published: (2026)
by: Samuel, Vinay, et al.
Published: (2026)
How to build trust in answers given by Generative AI for specific, and vague, financial questions
by: Zarifis, Alex, et al.
Published: (2024)
by: Zarifis, Alex, et al.
Published: (2024)
Multi-step retrieval and reasoning improves radiology question answering with large language models
by: Wind, Sebastian, et al.
Published: (2025)
by: Wind, Sebastian, et al.
Published: (2025)
Enhancing textual textbook question answering with large language models and retrieval augmented generation
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
by: Alawwad, Hessa Abdulrahman, et al.
Published: (2024)
Assessing the utility of a natural language processing model in answering common urological questions
by: Wyatt MacNevin, et al.
Published: (2025)
by: Wyatt MacNevin, et al.
Published: (2025)
Exploring Design Choices for Building Language-Specific LLMs
by: Tejaswi, Atula, et al.
Published: (2024)
by: Tejaswi, Atula, et al.
Published: (2024)
An $ab\;initio$ answer to long-debated questions about superconducting Nb$_3$Sn
by: Cucciari, Alessio, et al.
Published: (2025)
by: Cucciari, Alessio, et al.
Published: (2025)
Chapter “What is contemporary Japanese Cinema?”. Questioning the answers, answering with questions
by: CALORIO, GIACOMO
Published: (2022)
by: CALORIO, GIACOMO
Published: (2022)
Spoken question answering for visual queries
by: Shabtay, Nimrod, et al.
Published: (2025)
by: Shabtay, Nimrod, et al.
Published: (2025)
If generative AI is the answer, what is the question?
by: Tewari, Ambuj
Published: (2025)
by: Tewari, Ambuj
Published: (2025)
Similar Items
-
One ruler to measure them all: Benchmarking multilingual long-context language models
by: Kim, Yekyung, et al.
Published: (2025) -
One Thousand and One Pairs: A "novel" challenge for long-context language models
by: Karpinska, Marzena, et al.
Published: (2024) -
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
by: Russell, Jenna, et al.
Published: (2025) -
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025) -
Understanding Retrieval Augmentation for Long-Form Question Answering
by: Chen, Hung-Ting, et al.
Published: (2023)