Demo: Statistically Significant Results On Biases and Errors of LLMs Do Not Guarantee Generalizable Results
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jonathan, Qiu, Haoling, Lasko, Jonathan, Karakos, Damianos, Yarmohammadi, Mahsa, Dredze, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
Do Text-to-Vis Benchmarks Test Real Use of Visualisations?
von: Nguyen, Hy, et al.
Veröffentlicht: (2024)
von: Nguyen, Hy, et al.
Veröffentlicht: (2024)
Blind Spots and Biases: Exploring the Role of Annotator Cognitive Biases in NLP
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
Getting into the Flow: Towards Better Type Error Messages for Constraint-Based Type Inference
von: Bhanuka, Ishan, et al.
Veröffentlicht: (2024)
von: Bhanuka, Ishan, et al.
Veröffentlicht: (2024)
Do LLMs Truly Benefit from Longer Context in Automatic Post-Editing?
von: Kim, Ahrii, et al.
Veröffentlicht: (2026)
von: Kim, Ahrii, et al.
Veröffentlicht: (2026)
Your Students Don't Use LLMs Like You Wish They Did
von: Kobler, Sebastian, et al.
Veröffentlicht: (2026)
von: Kobler, Sebastian, et al.
Veröffentlicht: (2026)
Cognitively Biased Users Interacting with Algorithmically Biased Results in Whole-Session Search on Debated Topics
von: Wang, Ben, et al.
Veröffentlicht: (2024)
von: Wang, Ben, et al.
Veröffentlicht: (2024)
Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback
von: Tan, Mei, et al.
Veröffentlicht: (2026)
von: Tan, Mei, et al.
Veröffentlicht: (2026)
Power Echoes: Investigating Moderation Biases in Online Power-Asymmetric Conflicts
von: Li, Yaqiong, et al.
Veröffentlicht: (2026)
von: Li, Yaqiong, et al.
Veröffentlicht: (2026)
Exploring Gender Biases in Language Patterns of Human-Conversational Agent Conversations
von: Liu, Weizi
Veröffentlicht: (2024)
von: Liu, Weizi
Veröffentlicht: (2024)
Detection and Positive Reconstruction of Cognitive Distortion sentences: Mandarin Dataset and Evaluation
von: Lin, Shuya, et al.
Veröffentlicht: (2024)
von: Lin, Shuya, et al.
Veröffentlicht: (2024)
Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages
von: Shahid, Farhana, et al.
Veröffentlicht: (2025)
von: Shahid, Farhana, et al.
Veröffentlicht: (2025)
Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-Playing
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2024)
Prompts Matter: Comparing ML/GAI Approaches for Generating Inductive Qualitative Coding Results
von: Chen, John, et al.
Veröffentlicht: (2024)
von: Chen, John, et al.
Veröffentlicht: (2024)
LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs
von: Wu, Tongshuang, et al.
Veröffentlicht: (2023)
von: Wu, Tongshuang, et al.
Veröffentlicht: (2023)
SQLucid: Grounding Natural Language Database Queries with Interactive Explanations
von: Tian, Yuan, et al.
Veröffentlicht: (2024)
von: Tian, Yuan, et al.
Veröffentlicht: (2024)
Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
von: Shen, Hua, et al.
Veröffentlicht: (2025)
von: Shen, Hua, et al.
Veröffentlicht: (2025)
Gender Biases in Error Mitigation by Voice Assistants
von: Mahmood, Amama, et al.
Veröffentlicht: (2023)
von: Mahmood, Amama, et al.
Veröffentlicht: (2023)
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
von: Ivey, Jonathan, et al.
Veröffentlicht: (2024)
von: Ivey, Jonathan, et al.
Veröffentlicht: (2024)
Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models
von: Tran, Son Quoc, et al.
Veröffentlicht: (2025)
von: Tran, Son Quoc, et al.
Veröffentlicht: (2025)
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
von: Martin-Boyle, Anna, et al.
Veröffentlicht: (2026)
von: Martin-Boyle, Anna, et al.
Veröffentlicht: (2026)
CALYPSO: LLMs as Dungeon Masters' Assistants
von: Zhu, Andrew, et al.
Veröffentlicht: (2023)
von: Zhu, Andrew, et al.
Veröffentlicht: (2023)
(Ir)rationality and Cognitive Biases in Large Language Models
von: Macmillan-Scott, Olivia, et al.
Veröffentlicht: (2024)
von: Macmillan-Scott, Olivia, et al.
Veröffentlicht: (2024)
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective
von: Tak, Ala N., et al.
Veröffentlicht: (2024)
von: Tak, Ala N., et al.
Veröffentlicht: (2024)
Robots in the Middle: Evaluating LLMs in Dispute Resolution
von: Tan, Jinzhe, et al.
Veröffentlicht: (2024)
von: Tan, Jinzhe, et al.
Veröffentlicht: (2024)
LLMs Get Lost In Multi-Turn Conversation
von: Laban, Philippe, et al.
Veröffentlicht: (2025)
von: Laban, Philippe, et al.
Veröffentlicht: (2025)
Pragmatics beyond humans: meaning, communication, and LLMs
von: Gvoždiak, Vít
Veröffentlicht: (2025)
von: Gvoždiak, Vít
Veröffentlicht: (2025)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
von: Calderon, Nitay, et al.
Veröffentlicht: (2025)
LLMs Corrupt Your Documents When You Delegate
von: Laban, Philippe, et al.
Veröffentlicht: (2026)
von: Laban, Philippe, et al.
Veröffentlicht: (2026)
LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations
von: Salgado, Henry, et al.
Veröffentlicht: (2026)
von: Salgado, Henry, et al.
Veröffentlicht: (2026)
Revisiting OPRO: The Limitations of Small-Scale LLMs as Optimizers
von: Zhang, Tuo, et al.
Veröffentlicht: (2024)
von: Zhang, Tuo, et al.
Veröffentlicht: (2024)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
von: Fröhling, Leon, et al.
Veröffentlicht: (2024)
von: Fröhling, Leon, et al.
Veröffentlicht: (2024)
Generating Educational Materials with Different Levels of Readability using LLMs
von: Huang, Chieh-Yang, et al.
Veröffentlicht: (2024)
von: Huang, Chieh-Yang, et al.
Veröffentlicht: (2024)
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
von: Inie, Nanna, et al.
Veröffentlicht: (2023)
von: Inie, Nanna, et al.
Veröffentlicht: (2023)
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2025)
von: Kim, Seon Gyeom, et al.
Veröffentlicht: (2025)
Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with People
von: Huang, Dun-Ming, et al.
Veröffentlicht: (2024)
von: Huang, Dun-Ming, et al.
Veröffentlicht: (2024)
Agentic AutoSurvey: Let LLMs Survey LLMs
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows
von: Balashov, Yuri, et al.
Veröffentlicht: (2026)
von: Balashov, Yuri, et al.
Veröffentlicht: (2026)
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?
von: Choi, Alexander S., et al.
Veröffentlicht: (2024)
von: Choi, Alexander S., et al.
Veröffentlicht: (2024)
Leveraging Small LLMs for Argument Mining in Education: Argument Component Identification, Classification, and Assessment
von: Favero, Lucile, et al.
Veröffentlicht: (2025)
von: Favero, Lucile, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
von: Liu, Naiming, et al.
Veröffentlicht: (2025) -
Do Text-to-Vis Benchmarks Test Real Use of Visualisations?
von: Nguyen, Hy, et al.
Veröffentlicht: (2024) -
Blind Spots and Biases: Exploring the Role of Annotator Cognitive Biases in NLP
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024) -
Getting into the Flow: Towards Better Type Error Messages for Constraint-Based Type Inference
von: Bhanuka, Ishan, et al.
Veröffentlicht: (2024) -
Do LLMs Truly Benefit from Longer Context in Automatic Post-Editing?
von: Kim, Ahrii, et al.
Veröffentlicht: (2026)