More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Shafiei, Mohammadamin, Saffari, Hamidreza, Moosavi, Nafise Sadat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiHoax: A Dataset of Multi-hop False-Premise Questions
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
by: Saffari, Hamidreza, et al.
Published: (2024)
by: Saffari, Hamidreza, et al.
Published: (2024)
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
by: Shafiei, Mohammadamin, et al.
Published: (2025)
by: Shafiei, Mohammadamin, et al.
Published: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
by: Pandya, Mugdha, et al.
Published: (2024)
by: Pandya, Mugdha, et al.
Published: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026)
by: Kennedy, Ian W., et al.
Published: (2026)
How to Leverage Digit Embeddings to Represent Numbers?
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
by: Sivakumar, Jasivan Alex, et al.
Published: (2024)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
by: Mi, Maggie, et al.
Published: (2025)
by: Mi, Maggie, et al.
Published: (2025)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
by: Liu, Yiqi, et al.
Published: (2023)
by: Liu, Yiqi, et al.
Published: (2023)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
by: Mi, Maggie, et al.
Published: (2024)
by: Mi, Maggie, et al.
Published: (2024)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
by: Pastorino, Valeria, et al.
Published: (2024)
by: Pastorino, Valeria, et al.
Published: (2024)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
by: Xue, Huiyin, et al.
Published: (2025)
by: Xue, Huiyin, et al.
Published: (2025)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
by: Permadi, Vynska Amalia, et al.
Published: (2026)
by: Permadi, Vynska Amalia, et al.
Published: (2026)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
by: Aghaebe, Favour Yahdii, et al.
Published: (2026)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
by: James, Joseph, et al.
Published: (2026)
by: James, Joseph, et al.
Published: (2026)
Exploring Gender Disparities in Automatic Speech Recognition Technology
by: ElGhazaly, Hend, et al.
Published: (2025)
by: ElGhazaly, Hend, et al.
Published: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
Can I introduce my boyfriend to my grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification
by: Saffari, Hamidreza, et al.
Published: (2024)
by: Saffari, Hamidreza, et al.
Published: (2024)
Sell More, Play Less: Benchmarking LLM Realistic Selling Skill
by: Su, Xuanbo, et al.
Published: (2026)
by: Su, Xuanbo, et al.
Published: (2026)
More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
by: Vu, Duc Anh, et al.
Published: (2025)
by: Vu, Duc Anh, et al.
Published: (2025)
LIMO: Less is More for Reasoning
by: Ye, Yixin, et al.
Published: (2025)
by: Ye, Yixin, et al.
Published: (2025)
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
Less Is More: Cognitive Load and the Single-Prompt Ceiling in LLM Mathematical Reasoning
by: Cazares, Manuel Israel
Published: (2026)
by: Cazares, Manuel Israel
Published: (2026)
When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning
by: Sui, Yi, et al.
Published: (2026)
by: Sui, Yi, et al.
Published: (2026)
Acting Less is Reasoning More! Teaching Model to Act Efficiently
by: Wang, Hongru, et al.
Published: (2025)
by: Wang, Hongru, et al.
Published: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
by: Zhang, Qianchi, et al.
Published: (2025)
by: Zhang, Qianchi, et al.
Published: (2025)
Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information
by: Zhang, Yongheng, et al.
Published: (2024)
by: Zhang, Yongheng, et al.
Published: (2024)
Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection
by: Ayoobi, Navid, et al.
Published: (2025)
by: Ayoobi, Navid, et al.
Published: (2025)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
by: Akasaka, Muku, et al.
Published: (2026)
by: Akasaka, Muku, et al.
Published: (2026)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
by: Wang, Yihan, et al.
Published: (2022)
by: Wang, Yihan, et al.
Published: (2022)
Pensez: Less Data, Better Reasoning -- Rethinking French LLM
by: Ha, Huy Hoang
Published: (2025)
by: Ha, Huy Hoang
Published: (2025)
Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention
by: Yang, Lijie, et al.
Published: (2025)
by: Yang, Lijie, et al.
Published: (2025)
Similar Items
-
MultiHoax: A Dataset of Multi-hop False-Premise Questions
by: Shafiei, Mohammadamin, et al.
Published: (2025) -
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
by: Saffari, Hamidreza, et al.
Published: (2024) -
Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
by: Shafiei, Mohammadamin, et al.
Published: (2025) -
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
by: Pandya, Mugdha, et al.
Published: (2024) -
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
by: Kennedy, Ian W., et al.
Published: (2026)