See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Yulong, Liu, Yang, Yan, Jianhao, Bai, Xuefeng, Zhong, Ming, Yang, Yinghao, Yang, Ziyi, Zhu, Chenguang, Zhang, Yue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
por: Yang, Haoming, et al.
Publicado: (2025)
por: Yang, Haoming, et al.
Publicado: (2025)
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs
por: Wu, Yang, et al.
Publicado: (2025)
por: Wu, Yang, et al.
Publicado: (2025)
Constituency Parsing using LLMs
por: Bai, Xuefeng, et al.
Publicado: (2023)
por: Bai, Xuefeng, et al.
Publicado: (2023)
GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels
por: Yan, Jianhao, et al.
Publicado: (2024)
por: Yan, Jianhao, et al.
Publicado: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
por: Piot, Paloma, et al.
Publicado: (2025)
por: Piot, Paloma, et al.
Publicado: (2025)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
por: Yan, Jianhao, et al.
Publicado: (2024)
por: Yan, Jianhao, et al.
Publicado: (2024)
RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
por: Yan, Jianhao, et al.
Publicado: (2025)
por: Yan, Jianhao, et al.
Publicado: (2025)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
por: Son, Guijin, et al.
Publicado: (2024)
por: Son, Guijin, et al.
Publicado: (2024)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
por: Kidder, William, et al.
Publicado: (2024)
por: Kidder, William, et al.
Publicado: (2024)
Probing Minimalist Phase Structure in LLMs: What Universal Dependencies Cannot Represent
por: Chen, Yuanhao, et al.
Publicado: (2026)
por: Chen, Yuanhao, et al.
Publicado: (2026)
ELICIT: LLM Augmentation via External In-Context Capability
por: Wang, Futing, et al.
Publicado: (2024)
por: Wang, Futing, et al.
Publicado: (2024)
Potential and Challenges of Model Editing for Social Debiasing
por: Yan, Jianhao, et al.
Publicado: (2024)
por: Yan, Jianhao, et al.
Publicado: (2024)
Word Matters: What Influences Domain Adaptation in Summarization?
por: Li, Yinghao, et al.
Publicado: (2024)
por: Li, Yinghao, et al.
Publicado: (2024)
What Have We Achieved on Non-autoregressive Translation?
por: Li, Yafu, et al.
Publicado: (2024)
por: Li, Yafu, et al.
Publicado: (2024)
The Potential of LLMs in Medical Education: Generating Questions and Answers for Qualification Exams
por: Zhu, Yunqi, et al.
Publicado: (2024)
por: Zhu, Yunqi, et al.
Publicado: (2024)
Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics
por: Maya, Cristian Espinal
Publicado: (2026)
por: Maya, Cristian Espinal
Publicado: (2026)
Synthesizing Text-to-SQL Data from Weak and Strong LLMs
por: Yang, Jiaxi, et al.
Publicado: (2024)
por: Yang, Jiaxi, et al.
Publicado: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
por: Molfese, Francesco Maria, et al.
Publicado: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
por: Zhang, Hongbin, et al.
Publicado: (2025)
por: Zhang, Hongbin, et al.
Publicado: (2025)
LLM-based Translation Inference with Iterative Bilingual Understanding
por: Chen, Andong, et al.
Publicado: (2024)
por: Chen, Andong, et al.
Publicado: (2024)
Detecting RLVR Training Data via Structural Convergence of Reasoning
por: Zhang, Hongbo, et al.
Publicado: (2026)
por: Zhang, Hongbo, et al.
Publicado: (2026)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
por: Shi, Chufan, et al.
Publicado: (2026)
por: Shi, Chufan, et al.
Publicado: (2026)
Measuring Vision-Language STEM Skills of Neural Models
por: Shen, Jianhao, et al.
Publicado: (2024)
por: Shen, Jianhao, et al.
Publicado: (2024)
Measuring Social Norms of Large Language Models
por: Yuan, Ye, et al.
Publicado: (2024)
por: Yuan, Ye, et al.
Publicado: (2024)
Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
por: Shen, Weizhou, et al.
Publicado: (2024)
por: Shen, Weizhou, et al.
Publicado: (2024)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
por: Liang, Xiao, et al.
Publicado: (2025)
por: Liang, Xiao, et al.
Publicado: (2025)
LLM-based Discriminative Reasoning for Knowledge Graph Question Answering
por: Xu, Mufan, et al.
Publicado: (2024)
por: Xu, Mufan, et al.
Publicado: (2024)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
por: He, Jinwen, et al.
Publicado: (2023)
por: He, Jinwen, et al.
Publicado: (2023)
Improving Weak-to-Strong Generalization with Reliability-Aware Alignment
por: Guo, Yue, et al.
Publicado: (2024)
por: Guo, Yue, et al.
Publicado: (2024)
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
por: Chen, Zhou, et al.
Publicado: (2025)
por: Chen, Zhou, et al.
Publicado: (2025)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
por: Liu, Zijun, et al.
Publicado: (2024)
por: Liu, Zijun, et al.
Publicado: (2024)
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
por: Shan, Liang, et al.
Publicado: (2025)
por: Shan, Liang, et al.
Publicado: (2025)
Why LLMs Cannot Think and How to Fix It
por: Jahrens, Marius, et al.
Publicado: (2025)
por: Jahrens, Marius, et al.
Publicado: (2025)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
por: Huang, Shulin, et al.
Publicado: (2025)
por: Huang, Shulin, et al.
Publicado: (2025)
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
por: Xuan, Richmond Sin Jing, et al.
Publicado: (2025)
por: Xuan, Richmond Sin Jing, et al.
Publicado: (2025)
Knowledge Fusion of Chat LLMs: A Preliminary Technical Report
por: Wan, Fanqi, et al.
Publicado: (2024)
por: Wan, Fanqi, et al.
Publicado: (2024)
SEMDR: A Semantic-Aware Dual Encoder Model for Legal Judgment Prediction with Legal Clue Tracing
por: Liu, Pengjie, et al.
Publicado: (2024)
por: Liu, Pengjie, et al.
Publicado: (2024)
IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
por: Maimon, Aviya, et al.
Publicado: (2025)
por: Maimon, Aviya, et al.
Publicado: (2025)
RulePrompt: Weakly Supervised Text Classification with Prompting PLMs and Self-Iterative Logical Rules
por: Li, Miaomiao, et al.
Publicado: (2024)
por: Li, Miaomiao, et al.
Publicado: (2024)
LinguaLIFT: An Effective Two-stage Instruction Tuning Framework for Low-Resource Language Reasoning
por: Zhang, Hongbin, et al.
Publicado: (2024)
por: Zhang, Hongbin, et al.
Publicado: (2024)
Ejemplares similares
-
Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
por: Yang, Haoming, et al.
Publicado: (2025) -
Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs
por: Wu, Yang, et al.
Publicado: (2025) -
Constituency Parsing using LLMs
por: Bai, Xuefeng, et al.
Publicado: (2023) -
GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels
por: Yan, Jianhao, et al.
Publicado: (2024) -
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
por: Piot, Paloma, et al.
Publicado: (2025)