From Understanding to Generation: An Efficient Shortcut for Evaluating Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hangya, Viktor, Küch, Fabian, Gold, Darina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
di: Hangya, Viktor, et al.
Pubblicazione: (2023)
di: Hangya, Viktor, et al.
Pubblicazione: (2023)
Style-Specific Neurons for Steering LLMs in Text Style Transfer
di: Lai, Wen, et al.
Pubblicazione: (2024)
di: Lai, Wen, et al.
Pubblicazione: (2024)
Extending Multilingual Machine Translation through Imitation Learning
di: Lai, Wen, et al.
Pubblicazione: (2023)
di: Lai, Wen, et al.
Pubblicazione: (2023)
LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
di: Mirza, Paramita, et al.
Pubblicazione: (2025)
di: Mirza, Paramita, et al.
Pubblicazione: (2025)
Hate Personified: Investigating the role of LLMs in content moderation
di: Masud, Sarah, et al.
Pubblicazione: (2024)
di: Masud, Sarah, et al.
Pubblicazione: (2024)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
di: Yuan, Yu, et al.
Pubblicazione: (2024)
di: Yuan, Yu, et al.
Pubblicazione: (2024)
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
di: Chi, Ziheng, et al.
Pubblicazione: (2025)
di: Chi, Ziheng, et al.
Pubblicazione: (2025)
RAGONITE: Iterative Retrieval on Induced Databases and Verbalized RDF for Conversational QA over KGs with RAG
di: Roy, Rishiraj Saha, et al.
Pubblicazione: (2024)
di: Roy, Rishiraj Saha, et al.
Pubblicazione: (2024)
Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities
di: Lehmann, Hans Hergen, et al.
Pubblicazione: (2025)
di: Lehmann, Hans Hergen, et al.
Pubblicazione: (2025)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
di: Enström, Daniel, et al.
Pubblicazione: (2024)
di: Enström, Daniel, et al.
Pubblicazione: (2024)
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
di: Sakib, Fardin Ahsan, et al.
Pubblicazione: (2025)
di: Sakib, Fardin Ahsan, et al.
Pubblicazione: (2025)
Pre-Training LLMs on a budget: A comparison of three optimizers
di: Schlotthauer, Joel, et al.
Pubblicazione: (2025)
di: Schlotthauer, Joel, et al.
Pubblicazione: (2025)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
di: Zhou, Yuqing, et al.
Pubblicazione: (2024)
di: Zhou, Yuqing, et al.
Pubblicazione: (2024)
Investigating Multi-Hop Factual Shortcuts in Knowledge Editing of Large Language Models
di: Ju, Tianjie, et al.
Pubblicazione: (2024)
di: Ju, Tianjie, et al.
Pubblicazione: (2024)
Evaluating, Understanding, and Improving Constrained Text Generation for Large Language Models
di: Chen, Xiang, et al.
Pubblicazione: (2023)
di: Chen, Xiang, et al.
Pubblicazione: (2023)
Break the Chain: Large Language Models Can be Shortcut Reasoners
di: Ding, Mengru, et al.
Pubblicazione: (2024)
di: Ding, Mengru, et al.
Pubblicazione: (2024)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
di: Permadi, Vynska Amalia, et al.
Pubblicazione: (2026)
di: Permadi, Vynska Amalia, et al.
Pubblicazione: (2026)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
di: Wu, Yulong, et al.
Pubblicazione: (2025)
di: Wu, Yulong, et al.
Pubblicazione: (2025)
Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis
di: Zhu, Kejian, et al.
Pubblicazione: (2025)
di: Zhu, Kejian, et al.
Pubblicazione: (2025)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
di: Bihani, Geetanjali, et al.
Pubblicazione: (2024)
ltzGLUE: Luxembourgish General Language Understanding Evaluation
di: Plum, Alistair, et al.
Pubblicazione: (2026)
di: Plum, Alistair, et al.
Pubblicazione: (2026)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts?
di: Yang, Sohee, et al.
Pubblicazione: (2024)
di: Yang, Sohee, et al.
Pubblicazione: (2024)
Not Eliminate but Aggregate: Post-Hoc Control over Mixture-of-Experts to Address Shortcut Shifts in Natural Language Understanding
di: Honda, Ukyo, et al.
Pubblicazione: (2024)
di: Honda, Ukyo, et al.
Pubblicazione: (2024)
Mitigating Shortcut Reasoning in Language Models: A Gradient-Aware Training Approach
di: Cao, Hongyu, et al.
Pubblicazione: (2026)
di: Cao, Hongyu, et al.
Pubblicazione: (2026)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
di: Eshuijs, Leon, et al.
Pubblicazione: (2025)
Evaluating Chinese Ambiguity Understanding in Large Language Models
di: Mo, Junwen, et al.
Pubblicazione: (2026)
di: Mo, Junwen, et al.
Pubblicazione: (2026)
Investigating the Influence of Prompt-Specific Shortcuts in AI Generated Text Detection
di: Park, Choonghyun, et al.
Pubblicazione: (2024)
di: Park, Choonghyun, et al.
Pubblicazione: (2024)
Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation
di: Chhun, Cyril, et al.
Pubblicazione: (2024)
di: Chhun, Cyril, et al.
Pubblicazione: (2024)
Aligning Translation-Specific Understanding to General Understanding in Large Language Models
di: Huang, Yichong, et al.
Pubblicazione: (2024)
di: Huang, Yichong, et al.
Pubblicazione: (2024)
The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts
di: Madhusudan, Sangmitra, et al.
Pubblicazione: (2025)
di: Madhusudan, Sangmitra, et al.
Pubblicazione: (2025)
System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
di: Wang, Xiaoqiang, et al.
Pubblicazione: (2025)
di: Wang, Xiaoqiang, et al.
Pubblicazione: (2025)
Evaluating Language Models for Efficient Code Generation
di: Liu, Jiawei, et al.
Pubblicazione: (2024)
di: Liu, Jiawei, et al.
Pubblicazione: (2024)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
di: Yan, Lecheng, et al.
Pubblicazione: (2026)
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
di: Bhatia, Mehar, et al.
Pubblicazione: (2024)
di: Bhatia, Mehar, et al.
Pubblicazione: (2024)
AutoIntent: AutoML for Text Classification
di: Alekseev, Ilya, et al.
Pubblicazione: (2025)
di: Alekseev, Ilya, et al.
Pubblicazione: (2025)
Evaluating Dialect Robustness of Language Models via Conversation Understanding
di: Srirag, Dipankar, et al.
Pubblicazione: (2024)
di: Srirag, Dipankar, et al.
Pubblicazione: (2024)
AC-EVAL: Evaluating Ancient Chinese Language Understanding in Large Language Models
di: Wei, Yuting, et al.
Pubblicazione: (2024)
di: Wei, Yuting, et al.
Pubblicazione: (2024)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
Understanding Syntactic Generalization in Structure-inducing Language Models
di: Arps, David, et al.
Pubblicazione: (2025)
di: Arps, David, et al.
Pubblicazione: (2025)
Documenti analoghi
-
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
di: Hangya, Viktor, et al.
Pubblicazione: (2023) -
Style-Specific Neurons for Steering LLMs in Text Style Transfer
di: Lai, Wen, et al.
Pubblicazione: (2024) -
Extending Multilingual Machine Translation through Imitation Learning
di: Lai, Wen, et al.
Pubblicazione: (2023) -
LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
di: Mirza, Paramita, et al.
Pubblicazione: (2025) -
Hate Personified: Investigating the role of LLMs in content moderation
di: Masud, Sarah, et al.
Pubblicazione: (2024)