Do LLMs exhibit human-like response biases? A case study in survey design
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tjuatja, Lindia, Chen, Valerie, Wu, Sherry Tongshuang, Talwalkar, Ameet, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
par: Tjuatja, Lindia, et autres
Publié: (2025)
par: Tjuatja, Lindia, et autres
Publié: (2025)
What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
par: Tjuatja, Lindia, et autres
Publié: (2024)
par: Tjuatja, Lindia, et autres
Publié: (2024)
CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models
par: Sheikh, Zaid, et autres
Publié: (2024)
par: Sheikh, Zaid, et autres
Publié: (2024)
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text
par: Ginn, Michael, et autres
Publié: (2024)
par: Ginn, Michael, et autres
Publié: (2024)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
par: Kolawole, Steven, et autres
Publié: (2024)
par: Kolawole, Steven, et autres
Publié: (2024)
What do Language Models Learn and When? The Implicit Curriculum Hypothesis
par: Liu, Emmy, et autres
Publié: (2026)
par: Liu, Emmy, et autres
Publié: (2026)
Massively Multilingual Joint Segmentation and Glossing
par: Ginn, Michael, et autres
Publié: (2026)
par: Ginn, Michael, et autres
Publié: (2026)
Comparing Developer and LLM Biases in Code Evaluation
par: Mittal, Aditya, et autres
Publié: (2026)
par: Mittal, Aditya, et autres
Publié: (2026)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
par: Chen, Valerie, et autres
Publié: (2025)
par: Chen, Valerie, et autres
Publié: (2025)
Wav2Gloss: Generating Interlinear Glossed Text from Speech
par: He, Taiqi, et autres
Publié: (2024)
par: He, Taiqi, et autres
Publié: (2024)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
par: Gandhi, Saumya, et autres
Publié: (2024)
par: Gandhi, Saumya, et autres
Publié: (2024)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
par: Zhao, Chenyang, et autres
Publié: (2024)
par: Zhao, Chenyang, et autres
Publié: (2024)
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
par: Liu, Emmy, et autres
Publié: (2025)
par: Liu, Emmy, et autres
Publié: (2025)
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
par: Yu, Rebecca, et autres
Publié: (2025)
par: Yu, Rebecca, et autres
Publié: (2025)
Checklists Are Better Than Reward Models For Aligning Language Models
par: Viswanathan, Vijay, et autres
Publié: (2025)
par: Viswanathan, Vijay, et autres
Publié: (2025)
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
par: Liu, Jiarui, et autres
Publié: (2025)
par: Liu, Jiarui, et autres
Publié: (2025)
How can we assess human-agent interactions? Case studies in software agent design
par: Chen, Valerie, et autres
Publié: (2025)
par: Chen, Valerie, et autres
Publié: (2025)
Synthetic Multimodal Question Generation
par: Wu, Ian, et autres
Publié: (2024)
par: Wu, Ian, et autres
Publié: (2024)
"I didn't Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration
par: Kim, Eunsu, et autres
Publié: (2026)
par: Kim, Eunsu, et autres
Publié: (2026)
Do LLMs produce texts with "human-like" lexical diversity?
par: Kendro, Kelly, et autres
Publié: (2025)
par: Kendro, Kelly, et autres
Publié: (2025)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
par: Fernandes, Patrick, et autres
Publié: (2025)
par: Fernandes, Patrick, et autres
Publié: (2025)
Do LLMs exhibit the same commonsense capabilities across languages?
par: Martínez-Murillo, Ivan, et autres
Publié: (2025)
par: Martínez-Murillo, Ivan, et autres
Publié: (2025)
Completion $\neq$ Collaboration: Scaling Collaborative Effort with Agents
par: Shen, Shannon Zejiang, et autres
Publié: (2025)
par: Shen, Shannon Zejiang, et autres
Publié: (2025)
Do LLMs write like humans? Variation in grammatical and rhetorical styles
par: Reinhart, Alex, et autres
Publié: (2024)
par: Reinhart, Alex, et autres
Publié: (2024)
Go-Browse: Training Web Agents with Structured Exploration
par: Gandhi, Apurva, et autres
Publié: (2025)
par: Gandhi, Apurva, et autres
Publié: (2025)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
par: Soni, Aditya Bharat, et autres
Publié: (2025)
par: Soni, Aditya Bharat, et autres
Publié: (2025)
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
par: Pan, Jane, et autres
Publié: (2025)
par: Pan, Jane, et autres
Publié: (2025)
Effective Strategies for Asynchronous Software Engineering Agents
par: Geng, Jiayi, et autres
Publié: (2026)
par: Geng, Jiayi, et autres
Publié: (2026)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
par: He, Keyu, et autres
Publié: (2026)
par: He, Keyu, et autres
Publié: (2026)
Prompt-MII: Meta-Learning Instruction Induction for LLMs
par: Xiao, Emily, et autres
Publié: (2025)
par: Xiao, Emily, et autres
Publié: (2025)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
par: Chern, Steffi, et autres
Publié: (2024)
par: Chern, Steffi, et autres
Publié: (2024)
Demystifying Long Chain-of-Thought Reasoning in LLMs
par: Yeo, Edward, et autres
Publié: (2025)
par: Yeo, Edward, et autres
Publié: (2025)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
par: Nyandwi, Jean de Dieu, et autres
Publié: (2025)
par: Nyandwi, Jean de Dieu, et autres
Publié: (2025)
Learn Hard Problems During RL with Reference Guided Fine-tuning
par: Wu, Yangzhen, et autres
Publié: (2026)
par: Wu, Yangzhen, et autres
Publié: (2026)
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
par: Shen, Junhong, et autres
Publié: (2024)
par: Shen, Junhong, et autres
Publié: (2024)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
par: Kaneko, Masahiro, et autres
Publié: (2023)
par: Kaneko, Masahiro, et autres
Publié: (2023)
An Incomplete Loop: Instruction Inference, Instruction Following, and In-context Learning in Language Models
par: Liu, Emmy, et autres
Publié: (2024)
par: Liu, Emmy, et autres
Publié: (2024)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
par: Song, Yueqi, et autres
Publié: (2024)
par: Song, Yueqi, et autres
Publié: (2024)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
par: Zhang, Charlie, et autres
Publié: (2025)
par: Zhang, Charlie, et autres
Publié: (2025)
Midtraining Bridges Pretraining and Posttraining Distributions
par: Liu, Emmy, et autres
Publié: (2025)
par: Liu, Emmy, et autres
Publié: (2025)
Documents similaires
-
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
par: Tjuatja, Lindia, et autres
Publié: (2025) -
What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length
par: Tjuatja, Lindia, et autres
Publié: (2024) -
CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models
par: Sheikh, Zaid, et autres
Publié: (2024) -
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text
par: Ginn, Michael, et autres
Publié: (2024) -
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
par: Kolawole, Steven, et autres
Publié: (2024)