See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yulong, Liu, Yang, Yan, Jianhao, Bai, Xuefeng, Zhong, Ming, Yang, Yinghao, Yang, Ziyi, Zhu, Chenguang, Zhang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
by: Yang, Haoming, et al.
Published: (2025)
by: Yang, Haoming, et al.
Published: (2025)
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs
by: Wu, Yang, et al.
Published: (2025)
by: Wu, Yang, et al.
Published: (2025)
Constituency Parsing using LLMs
by: Bai, Xuefeng, et al.
Published: (2023)
by: Bai, Xuefeng, et al.
Published: (2023)
GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
by: Kidder, William, et al.
Published: (2024)
by: Kidder, William, et al.
Published: (2024)
Probing Minimalist Phase Structure in LLMs: What Universal Dependencies Cannot Represent
by: Chen, Yuanhao, et al.
Published: (2026)
by: Chen, Yuanhao, et al.
Published: (2026)
ELICIT: LLM Augmentation via External In-Context Capability
by: Wang, Futing, et al.
Published: (2024)
by: Wang, Futing, et al.
Published: (2024)
Potential and Challenges of Model Editing for Social Debiasing
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
Word Matters: What Influences Domain Adaptation in Summarization?
by: Li, Yinghao, et al.
Published: (2024)
by: Li, Yinghao, et al.
Published: (2024)
What Have We Achieved on Non-autoregressive Translation?
by: Li, Yafu, et al.
Published: (2024)
by: Li, Yafu, et al.
Published: (2024)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
by: Zhu, Yunqi, et al.
Published: (2024)
by: Zhu, Yunqi, et al.
Published: (2024)
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
by: Maya, Cristian Espinal
Published: (2026)
by: Maya, Cristian Espinal
Published: (2026)
Synthesizing Text-to-SQL Data from Weak and Strong LLMs
by: Yang, Jiaxi, et al.
Published: (2024)
by: Yang, Jiaxi, et al.
Published: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
LLM-based Translation Inference with Iterative Bilingual Understanding
by: Chen, Andong, et al.
Published: (2024)
by: Chen, Andong, et al.
Published: (2024)
Detecting RLVR Training Data via Structural Convergence of Reasoning
by: Zhang, Hongbo, et al.
Published: (2026)
by: Zhang, Hongbo, et al.
Published: (2026)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
Measuring Vision-Language STEM Skills of Neural Models
by: Shen, Jianhao, et al.
Published: (2024)
by: Shen, Jianhao, et al.
Published: (2024)
Measuring Social Norms of Large Language Models
by: Yuan, Ye, et al.
Published: (2024)
by: Yuan, Ye, et al.
Published: (2024)
Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
by: Shen, Weizhou, et al.
Published: (2024)
by: Shen, Weizhou, et al.
Published: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
by: Liang, Xiao, et al.
Published: (2025)
by: Liang, Xiao, et al.
Published: (2025)
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
by: Xu, Mufan, et al.
Published: (2024)
by: Xu, Mufan, et al.
Published: (2024)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
by: He, Jinwen, et al.
Published: (2023)
by: He, Jinwen, et al.
Published: (2023)
Improving Weak-to-Strong Generalization with Reliability-Aware Alignment
by: Guo, Yue, et al.
Published: (2024)
by: Guo, Yue, et al.
Published: (2024)
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
by: Liu, Zijun, et al.
Published: (2024)
by: Liu, Zijun, et al.
Published: (2024)
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
by: Shan, Liang, et al.
Published: (2025)
by: Shan, Liang, et al.
Published: (2025)
Why LLMs Cannot Think and How to Fix It
by: Jahrens, Marius, et al.
Published: (2025)
by: Jahrens, Marius, et al.
Published: (2025)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
by: Huang, Shulin, et al.
Published: (2025)
by: Huang, Shulin, et al.
Published: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
by: Xuan, Richmond Sin Jing, et al.
Published: (2025)
by: Xuan, Richmond Sin Jing, et al.
Published: (2025)
Knowledge Fusion of Chat LLMs: A Preliminary Technical Report
by: Wan, Fanqi, et al.
Published: (2024)
by: Wan, Fanqi, et al.
Published: (2024)
SEMDR: A Semantic-Aware Dual Encoder Model for Legal Judgment Prediction with Legal Clue Tracing
by: Liu, Pengjie, et al.
Published: (2024)
by: Liu, Pengjie, et al.
Published: (2024)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
by: Maimon, Aviya, et al.
Published: (2025)
by: Maimon, Aviya, et al.
Published: (2025)
RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical Rules
by: Li, Miaomiao, et al.
Published: (2024)
by: Li, Miaomiao, et al.
Published: (2024)
LinguaLIFT: An Effective Two-stage Instruction Tuning Framework for Low-Resource Language Reasoning
by: Zhang, Hongbin, et al.
Published: (2024)
by: Zhang, Hongbin, et al.
Published: (2024)
Similar Items
-
Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
by: Yang, Haoming, et al.
Published: (2025) -
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs
by: Wu, Yang, et al.
Published: (2025) -
Constituency Parsing using LLMs
by: Bai, Xuefeng, et al.
Published: (2023) -
GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels
by: Yan, Jianhao, et al.
Published: (2024) -
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)