Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Chochlakis, Georgios, Pandiyan, Niyantha Maruthu, Lerman, Kristina, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models Do Multi-Label Classification Differently
by: Ma, Marcus, et al.
Published: (2025)
by: Ma, Marcus, et al.
Published: (2025)
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
by: Chochlakis, Georgios, et al.
Published: (2025)
by: Chochlakis, Georgios, et al.
Published: (2025)
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Authors Should Label Their Own Documents
by: Ma, Marcus, et al.
Published: (2025)
by: Ma, Marcus, et al.
Published: (2025)
Don't Blame the Data, Blame the Model: Understanding Noise and Bias When Learning from Subjective Annotations
by: Anand, Abhishek, et al.
Published: (2024)
by: Anand, Abhishek, et al.
Published: (2024)
Why Chain of Thought Fails in Clinical Text Understanding
by: Wu, Jiageng, et al.
Published: (2025)
by: Wu, Jiageng, et al.
Published: (2025)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
by: Khan, Imran
Published: (2025)
by: Khan, Imran
Published: (2025)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Think, But Don't Overthink: Reproducing Recursive Language Models
by: Wang, Daren
Published: (2026)
by: Wang, Daren
Published: (2026)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
by: Hassid, Michael, et al.
Published: (2025)
by: Hassid, Michael, et al.
Published: (2025)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
by: Goloburda, Maiya, et al.
Published: (2026)
by: Goloburda, Maiya, et al.
Published: (2026)
How to Retrieve Examples in In-context Learning to Improve Conversational Emotion Recognition using Large Language Models?
by: Wang, Mengqi, et al.
Published: (2025)
by: Wang, Mengqi, et al.
Published: (2025)
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
by: Hernandez, Adriano
Published: (2024)
by: Hernandez, Adriano
Published: (2024)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
by: Yan, Shaotian, et al.
Published: (2025)
by: Yan, Shaotian, et al.
Published: (2025)
Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
by: Lumer, Elias, et al.
Published: (2026)
by: Lumer, Elias, et al.
Published: (2026)
Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks
by: Mokhberian, Negar, et al.
Published: (2023)
by: Mokhberian, Negar, et al.
Published: (2023)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
by: Chan, Brian J, et al.
Published: (2024)
by: Chan, Brian J, et al.
Published: (2024)
Why Don't You Read This?
by: Chesley, Robert E.
Published: (1970)
by: Chesley, Robert E.
Published: (1970)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
by: Tyukin, Georgy, et al.
Published: (2024)
by: Tyukin, Georgy, et al.
Published: (2024)
Don't Erase, Inform! Detecting and Contextualizing Harmful Language in Cultural Heritage Collections
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
by: Mastromichalakis, Orfeas Menis, et al.
Published: (2025)
CHATTER: A Character Attribution Dataset for Narrative Understanding
by: Baruah, Sabyasachee, et al.
Published: (2024)
by: Baruah, Sabyasachee, et al.
Published: (2024)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
Intelligence Requires Grounding But Not Embodiment
by: Ma, Marcus, et al.
Published: (2026)
by: Ma, Marcus, et al.
Published: (2026)
Fine-Tune, Don't Prompt, Your Language Model to Identify Biased Language in Clinical Notes
by: Landi, Isotta, et al.
Published: (2026)
by: Landi, Isotta, et al.
Published: (2026)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
by: Yu, Zhiyuan, et al.
Published: (2024)
by: Yu, Zhiyuan, et al.
Published: (2024)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
Active Prompting with Chain-of-Thought for Large Language Models
by: Diao, Shizhe, et al.
Published: (2023)
by: Diao, Shizhe, et al.
Published: (2023)
Semantic F1 Scores: Fair Evaluation Under Fuzzy Class Boundaries
by: Chochlakis, Georgios, et al.
Published: (2025)
by: Chochlakis, Georgios, et al.
Published: (2025)
CPL-NoViD: Context-Aware Prompt-based Learning for Norm Violation Detection in Online Communities
by: He, Zihao, et al.
Published: (2023)
by: He, Zihao, et al.
Published: (2023)
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
by: Pakhomov, Egor, et al.
Published: (2025)
by: Pakhomov, Egor, et al.
Published: (2025)
Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems
by: Feng, Tao, et al.
Published: (2026)
by: Feng, Tao, et al.
Published: (2026)
ChainLM: Empowering Large Language Models with Improved Chain-of-Thought Prompting
by: Cheng, Xiaoxue, et al.
Published: (2024)
by: Cheng, Xiaoxue, et al.
Published: (2024)
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
by: Li, Xiaomin, et al.
Published: (2025)
by: Li, Xiaomin, et al.
Published: (2025)
Why Can Large Language Models Generate Correct Chain-of-Thoughts?
by: Tutunov, Rasul, et al.
Published: (2023)
by: Tutunov, Rasul, et al.
Published: (2023)
Similar Items
-
Large Language Models Do Multi-Label Classification Differently
by: Ma, Marcus, et al.
Published: (2025) -
Aggregation Artifacts in Subjective Tasks Collapse Large Language Models' Posteriors
by: Chochlakis, Georgios, et al.
Published: (2024) -
Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
by: Chochlakis, Georgios, et al.
Published: (2025) -
The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
by: Chochlakis, Georgios, et al.
Published: (2024) -
Authors Should Label Their Own Documents
by: Ma, Marcus, et al.
Published: (2025)