Gespeichert in:
| Hauptverfasser: | Srikanth, Neha, Carpuat, Marine, Rudinger, Rachel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.11717 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
von: Srikanth, Neha, et al.
Veröffentlicht: (2025)
Multiple LLM Agents Debate for Equitable Cultural Alignment
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
von: Acquaye, Christabel, et al.
Veröffentlicht: (2026)
von: Acquaye, Christabel, et al.
Veröffentlicht: (2026)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2026)
von: Srikanth, Neha, et al.
Veröffentlicht: (2026)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
von: Palta, Shramay, et al.
Veröffentlicht: (2024)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
How Multilingual Are Large Language Models Fine-Tuned for Translation?
von: Richburg, Aquia, et al.
Veröffentlicht: (2024)
von: Richburg, Aquia, et al.
Veröffentlicht: (2024)
Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat Logs
von: Sarkar, Rupak, et al.
Veröffentlicht: (2025)
von: Sarkar, Rupak, et al.
Veröffentlicht: (2025)
Steering Large Language Models with Register Analysis for Arbitrary Style Transfer
von: Yang, Xinchen, et al.
Veröffentlicht: (2025)
von: Yang, Xinchen, et al.
Veröffentlicht: (2025)
Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension
von: Agrawal, Sweta, et al.
Veröffentlicht: (2023)
von: Agrawal, Sweta, et al.
Veröffentlicht: (2023)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
von: Ki, Dayeon, et al.
Veröffentlicht: (2024)
von: Ki, Dayeon, et al.
Veröffentlicht: (2024)
Keep It Private: Unsupervised Privatization of Online Text
von: Bao, Calvin, et al.
Veröffentlicht: (2024)
von: Bao, Calvin, et al.
Veröffentlicht: (2024)
Should We be Pedantic About Reasoning Errors in Machine Translation?
von: Bao, Calvin, et al.
Veröffentlicht: (2026)
von: Bao, Calvin, et al.
Veröffentlicht: (2026)
Automatic Input Rewriting Improves Translation with Large Language Models
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
AskQE: Question Answering as Automatic Evaluation for Machine Translation
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2023)
von: Srikanth, Neha, et al.
Veröffentlicht: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
What Makes Good Multilingual Reasoning? Disentangling Reasoning Traces with Measurable Features
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
SpeechQE: Estimating the Quality of Direct Speech Translation
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
von: Balepur, Nishant, et al.
Veröffentlicht: (2024)
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
von: Ravisankar, Kartik, et al.
Veröffentlicht: (2025)
von: Ravisankar, Kartik, et al.
Veröffentlicht: (2025)
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S
von: Acquaye, Christabel, et al.
Veröffentlicht: (2024)
von: Acquaye, Christabel, et al.
Veröffentlicht: (2024)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
von: Sancheti, Abhilasha, et al.
Veröffentlicht: (2024)
von: Sancheti, Abhilasha, et al.
Veröffentlicht: (2024)
Multilingual large language models leak human stereotypes across language boundaries
von: Cao, Yang Trista, et al.
Veröffentlicht: (2023)
von: Cao, Yang Trista, et al.
Veröffentlicht: (2023)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
von: Balepur, Nishant, et al.
Veröffentlicht: (2023)
von: Balepur, Nishant, et al.
Veröffentlicht: (2023)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
von: Zhao, Lingjun, et al.
Veröffentlicht: (2026)
von: Zhao, Lingjun, et al.
Veröffentlicht: (2026)
Words as Bridges: Exploring Computational Support for Cross-Disciplinary Translation Work
von: Bao, Calvin, et al.
Veröffentlicht: (2025)
von: Bao, Calvin, et al.
Veröffentlicht: (2025)
On the Mutual Influence of Gender and Occupation in LLM Representations
von: An, Haozhe, et al.
Veröffentlicht: (2025)
von: An, Haozhe, et al.
Veröffentlicht: (2025)
Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations
von: Xiao, Yimin, et al.
Veröffentlicht: (2025)
von: Xiao, Yimin, et al.
Veröffentlicht: (2025)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
von: Balepur, Nishant, et al.
Veröffentlicht: (2025)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
von: Hou, Yu, et al.
Veröffentlicht: (2025)
von: Hou, Yu, et al.
Veröffentlicht: (2025)
'Rich Dad, Poor Lad': How do Large Language Models Contextualize Socioeconomic Factors in College Admission ?
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
von: Baumler, Connor, et al.
Veröffentlicht: (2026)
von: Baumler, Connor, et al.
Veröffentlicht: (2026)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
Do Large Language Models Discriminate in Hiring Decisions on the Basis of Race, Ethnicity, and Gender?
von: An, Haozhe, et al.
Veröffentlicht: (2024)
von: An, Haozhe, et al.
Veröffentlicht: (2024)
Speaking the Right Language: The Impact of Expertise Alignment in User-AI Interactions
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
von: Palta, Shramay, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
von: Srikanth, Neha, et al.
Veröffentlicht: (2025) -
Multiple LLM Agents Debate for Equitable Cultural Alignment
von: Ki, Dayeon, et al.
Veröffentlicht: (2025) -
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
von: Acquaye, Christabel, et al.
Veröffentlicht: (2026) -
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
von: Srikanth, Neha, et al.
Veröffentlicht: (2026) -
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
von: Palta, Shramay, et al.
Veröffentlicht: (2024)