"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Xinpeng, Ma, Bolei, Hu, Chengzhi, Weber-Genzel, Leon, Röttger, Paul, Kreuter, Frauke, Hovy, Dirk, Plank, Barbara |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
di: Wang, Xinpeng, et al.
Pubblicazione: (2024)
di: Wang, Xinpeng, et al.
Pubblicazione: (2024)
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
di: Wang, Xinpeng, et al.
Pubblicazione: (2024)
di: Wang, Xinpeng, et al.
Pubblicazione: (2024)
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2023)
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2023)
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
di: Ma, Bolei, et al.
Pubblicazione: (2024)
di: Ma, Bolei, et al.
Pubblicazione: (2024)
Position: Insights from Survey Methodology can Improve Training Data
di: Eckman, Stephanie, et al.
Pubblicazione: (2024)
di: Eckman, Stephanie, et al.
Pubblicazione: (2024)
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
di: Eckman, Stephanie, et al.
Pubblicazione: (2025)
di: Eckman, Stephanie, et al.
Pubblicazione: (2025)
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study
di: Ma, Bolei, et al.
Pubblicazione: (2024)
di: Ma, Bolei, et al.
Pubblicazione: (2024)
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
di: Rooein, Donya, et al.
Pubblicazione: (2024)
di: Rooein, Donya, et al.
Pubblicazione: (2024)
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
di: Pernisi, Fabio, et al.
Pubblicazione: (2024)
di: Pernisi, Fabio, et al.
Pubblicazione: (2024)
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
di: Ma, Bolei, et al.
Pubblicazione: (2025)
di: Ma, Bolei, et al.
Pubblicazione: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
di: Röttger, Paul, et al.
Pubblicazione: (2024)
di: Röttger, Paul, et al.
Pubblicazione: (2024)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
di: Orlikowski, Matthias, et al.
Pubblicazione: (2023)
di: Orlikowski, Matthias, et al.
Pubblicazione: (2023)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
di: Russo, Giuseppe, et al.
Pubblicazione: (2025)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
di: Baan, Joris, et al.
Pubblicazione: (2026)
di: Baan, Joris, et al.
Pubblicazione: (2026)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
di: Orlikowski, Matthias, et al.
Pubblicazione: (2025)
di: Orlikowski, Matthias, et al.
Pubblicazione: (2025)
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks
di: Ma, Bolei, et al.
Pubblicazione: (2024)
di: Ma, Bolei, et al.
Pubblicazione: (2024)
The Call for Socially Aware Language Technologies
di: Yang, Diyi, et al.
Pubblicazione: (2024)
di: Yang, Diyi, et al.
Pubblicazione: (2024)
VariErr NLI: Separating Annotation Error from Human Label Variation
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2024)
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2024)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance
di: de Araujo, Pedro Henrique Luz, et al.
Pubblicazione: (2025)
di: de Araujo, Pedro Henrique Luz, et al.
Pubblicazione: (2025)
Do Clickers Improve Library Instruction? Lock in Your Answers Now
di: Dill, Emily
Pubblicazione: (2008)
di: Dill, Emily
Pubblicazione: (2008)
Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting
di: Ball, Sarah, et al.
Pubblicazione: (2026)
di: Ball, Sarah, et al.
Pubblicazione: (2026)
Rehearsing Answers to Probable Questions with Perspective-Taking
di: Shih, Yung-Yu, et al.
Pubblicazione: (2024)
di: Shih, Yung-Yu, et al.
Pubblicazione: (2024)
Chapter 7 Digital Trace Data
di: Keusch, Florian, et al.
Pubblicazione: (2021)
di: Keusch, Florian, et al.
Pubblicazione: (2021)
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
di: Ma, Bolei, et al.
Pubblicazione: (2025)
di: Ma, Bolei, et al.
Pubblicazione: (2025)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
di: Yuan, Chenchen, et al.
Pubblicazione: (2026)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
di: Kern, Christoph, et al.
Pubblicazione: (2023)
di: Kern, Christoph, et al.
Pubblicazione: (2023)
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2023)
di: Plaza-del-Arco, Flor Miriam, et al.
Pubblicazione: (2023)
Diffusion Language Models Are Natively Length-Aware
di: Rossi, Vittorio, et al.
Pubblicazione: (2026)
di: Rossi, Vittorio, et al.
Pubblicazione: (2026)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
di: Shi, Xiaofeng, et al.
Pubblicazione: (2025)
di: Shi, Xiaofeng, et al.
Pubblicazione: (2025)
RefChartQA: Grounding Visual Answer on Chart Images through Instruction Tuning
di: Vogel, Alexander, et al.
Pubblicazione: (2025)
di: Vogel, Alexander, et al.
Pubblicazione: (2025)
Two-Stage Quranic QA via Ensemble Retrieval and Instruction-Tuned Answer Extraction
di: Basem, Mohamed, et al.
Pubblicazione: (2025)
di: Basem, Mohamed, et al.
Pubblicazione: (2025)
Instruction Data Selection via Answer Divergence
di: Li, Bo, et al.
Pubblicazione: (2026)
di: Li, Bo, et al.
Pubblicazione: (2026)
How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning
di: Chen, Haoyang, et al.
Pubblicazione: (2026)
di: Chen, Haoyang, et al.
Pubblicazione: (2026)
Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
di: Gonzalez-Gutierrez, Cesar, et al.
Pubblicazione: (2025)
di: Gonzalez-Gutierrez, Cesar, et al.
Pubblicazione: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
di: Röttger, Paul, et al.
Pubblicazione: (2023)
di: Röttger, Paul, et al.
Pubblicazione: (2023)
Approaches to Text Summarization: Questions and Answers
di: Salvador Climent
Pubblicazione: (2004)
di: Salvador Climent
Pubblicazione: (2004)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
di: Baumann, Joachim, et al.
Pubblicazione: (2025)
Answer is All You Need: Instruction-following Text Embedding via Answering the Question
di: Peng, Letian, et al.
Pubblicazione: (2024)
di: Peng, Letian, et al.
Pubblicazione: (2024)
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
di: Chen, Zhou, et al.
Pubblicazione: (2025)
di: Chen, Zhou, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
di: Wang, Xinpeng, et al.
Pubblicazione: (2024) -
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
di: Wang, Xinpeng, et al.
Pubblicazione: (2024) -
Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?
di: Weber-Genzel, Leon, et al.
Pubblicazione: (2023) -
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models
di: Ma, Bolei, et al.
Pubblicazione: (2024) -
Position: Insights from Survey Methodology can Improve Training Data
di: Eckman, Stephanie, et al.
Pubblicazione: (2024)