From Understanding to Generation: An Efficient Shortcut for Evaluating Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hangya, Viktor, Küch, Fabian, Gold, Darina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
by: Hangya, Viktor, et al.
Published: (2023)
by: Hangya, Viktor, et al.
Published: (2023)
Style-Specific Neurons for Steering LLMs in Text Style Transfer
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
Extending Multilingual Machine Translation through Imitation Learning
by: Lai, Wen, et al.
Published: (2023)
by: Lai, Wen, et al.
Published: (2023)
LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
by: Mirza, Paramita, et al.
Published: (2025)
by: Mirza, Paramita, et al.
Published: (2025)
Hate Personified: Investigating the role of LLMs in content moderation
by: Masud, Sarah, et al.
Published: (2024)
by: Masud, Sarah, et al.
Published: (2024)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
by: Yuan, Yu, et al.
Published: (2024)
by: Yuan, Yu, et al.
Published: (2024)
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
by: Chi, Ziheng, et al.
Published: (2025)
by: Chi, Ziheng, et al.
Published: (2025)
RAGONITE: Iterative Retrieval on Induced Databases and Verbalized RDF for Conversational QA over KGs with RAG
by: Roy, Rishiraj Saha, et al.
Published: (2024)
by: Roy, Rishiraj Saha, et al.
Published: (2024)
Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
by: Lehmann, Hans Hergen, et al.
Published: (2025)
by: Lehmann, Hans Hergen, et al.
Published: (2025)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)
by: Enström, Daniel, et al.
Published: (2024)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
by: Sakib, Fardin Ahsan, et al.
Published: (2025)
Pre-Training LLMs on a budget: A comparison of three optimizers
by: Schlotthauer, Joel, et al.
Published: (2025)
by: Schlotthauer, Joel, et al.
Published: (2025)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
by: Zhou, Yuqing, et al.
Published: (2024)
by: Zhou, Yuqing, et al.
Published: (2024)
Investigating Multi-Hop Factual Shortcuts in Knowledge Editing of Large Language Models
by: Ju, Tianjie, et al.
Published: (2024)
by: Ju, Tianjie, et al.
Published: (2024)
Evaluating, Understanding, and Improving Constrained Text Generation for Large Language Models
by: Chen, Xiang, et al.
Published: (2023)
by: Chen, Xiang, et al.
Published: (2023)
Break the Chain: Large Language Models Can be Shortcut Reasoners
by: Ding, Mengru, et al.
Published: (2024)
by: Ding, Mengru, et al.
Published: (2024)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
by: Permadi, Vynska Amalia, et al.
Published: (2026)
by: Permadi, Vynska Amalia, et al.
Published: (2026)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
by: Wu, Yulong, et al.
Published: (2025)
by: Wu, Yulong, et al.
Published: (2025)
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
by: Zhu, Kejian, et al.
Published: (2025)
by: Zhu, Kejian, et al.
Published: (2025)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
by: Bihani, Geetanjali, et al.
Published: (2024)
by: Bihani, Geetanjali, et al.
Published: (2024)
ltzGLUE: Luxembourgish General Language Understanding Evaluation
by: Plum, Alistair, et al.
Published: (2026)
by: Plum, Alistair, et al.
Published: (2026)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
by: Yang, Sohee, et al.
Published: (2024)
by: Yang, Sohee, et al.
Published: (2024)
Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language Understanding
by: Honda, Ukyo, et al.
Published: (2024)
by: Honda, Ukyo, et al.
Published: (2024)
Mitigating Shortcut Reasoning in Language Models: A Gradient-Aware Training Approach
by: Cao, Hongyu, et al.
Published: (2026)
by: Cao, Hongyu, et al.
Published: (2026)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Evaluating Chinese Ambiguity Understanding in Large Language Models
by: Mo, Junwen, et al.
Published: (2026)
by: Mo, Junwen, et al.
Published: (2026)
Investigating the Influence of Prompt-Specific Shortcuts in AI Generated Text Detection
by: Park, Choonghyun, et al.
Published: (2024)
by: Park, Choonghyun, et al.
Published: (2024)
Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation
by: Chhun, Cyril, et al.
Published: (2024)
by: Chhun, Cyril, et al.
Published: (2024)
Aligning Translation-Specific Understanding to General Understanding in Large Language Models
by: Huang, Yichong, et al.
Published: (2024)
by: Huang, Yichong, et al.
Published: (2024)
The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
by: Madhusudan, Sangmitra, et al.
Published: (2025)
by: Madhusudan, Sangmitra, et al.
Published: (2025)
System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
by: Wang, Xiaoqiang, et al.
Published: (2025)
by: Wang, Xiaoqiang, et al.
Published: (2025)
Evaluating Language Models for Efficient Code Generation
by: Liu, Jiawei, et al.
Published: (2024)
by: Liu, Jiawei, et al.
Published: (2024)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
by: Bhatia, Mehar, et al.
Published: (2024)
by: Bhatia, Mehar, et al.
Published: (2024)
AutoIntent: AutoML for Text Classification
by: Alekseev, Ilya, et al.
Published: (2025)
by: Alekseev, Ilya, et al.
Published: (2025)
Evaluating Dialect Robustness of Language Models via Conversation Understanding
by: Srirag, Dipankar, et al.
Published: (2024)
by: Srirag, Dipankar, et al.
Published: (2024)
AC-EVAL: Evaluating Ancient Chinese Language Understanding in Large Language Models
by: Wei, Yuting, et al.
Published: (2024)
by: Wei, Yuting, et al.
Published: (2024)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Understanding Syntactic Generalization in Structure-inducing Language Models
by: Arps, David, et al.
Published: (2025)
by: Arps, David, et al.
Published: (2025)
Similar Items
-
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
by: Hangya, Viktor, et al.
Published: (2023) -
Style-Specific Neurons for Steering LLMs in Text Style Transfer
by: Lai, Wen, et al.
Published: (2024) -
Extending Multilingual Machine Translation through Imitation Learning
by: Lai, Wen, et al.
Published: (2023) -
LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
by: Mirza, Paramita, et al.
Published: (2025) -
Hate Personified: Investigating the role of LLMs in content moderation
by: Masud, Sarah, et al.
Published: (2024)