The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zihao, Jiang, Yibo, Yu, Jiahao, Huang, Heqing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Pun Unintended: LLMs and the Illusion of Humor Understanding
di: Zangari, Alessandro, et al.
Pubblicazione: (2025)
di: Zangari, Alessandro, et al.
Pubblicazione: (2025)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
di: Zong, Chang, et al.
Pubblicazione: (2024)
di: Zong, Chang, et al.
Pubblicazione: (2024)
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
The Impact of Role Design in In-Context Learning for Large Language Models
di: Rouzegar, Hamidreza, et al.
Pubblicazione: (2025)
di: Rouzegar, Hamidreza, et al.
Pubblicazione: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
di: Yang, Yibo
Pubblicazione: (2025)
di: Yang, Yibo
Pubblicazione: (2025)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
di: Štefánik, Michal, et al.
Pubblicazione: (2025)
di: Štefánik, Michal, et al.
Pubblicazione: (2025)
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
di: Wang, Yongjie, et al.
Pubblicazione: (2025)
Transforming and Combining Rewards for Aligning Large Language Models
di: Wang, Zihao, et al.
Pubblicazione: (2024)
di: Wang, Zihao, et al.
Pubblicazione: (2024)
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
di: Tu, Songjun, et al.
Pubblicazione: (2025)
di: Tu, Songjun, et al.
Pubblicazione: (2025)
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
di: Takahashi, Kosuke, et al.
Pubblicazione: (2024)
di: Takahashi, Kosuke, et al.
Pubblicazione: (2024)
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning
di: Ming, Xiaoyang, et al.
Pubblicazione: (2026)
di: Ming, Xiaoyang, et al.
Pubblicazione: (2026)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation
di: Xu, Zhe, et al.
Pubblicazione: (2024)
di: Xu, Zhe, et al.
Pubblicazione: (2024)
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
di: Huang, Donghao, et al.
Pubblicazione: (2026)
di: Huang, Donghao, et al.
Pubblicazione: (2026)
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
di: Fang, Xi, et al.
Pubblicazione: (2025)
di: Fang, Xi, et al.
Pubblicazione: (2025)
Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts
di: Xu, Beining, et al.
Pubblicazione: (2025)
di: Xu, Beining, et al.
Pubblicazione: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
Shortcuts Arising from Contrast: Effective and Covert Clean-Label Attacks in Prompt-Based Learning
di: Xie, Xiaopeng, et al.
Pubblicazione: (2024)
di: Xie, Xiaopeng, et al.
Pubblicazione: (2024)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
di: Mehta, Rahul, et al.
Pubblicazione: (2024)
di: Mehta, Rahul, et al.
Pubblicazione: (2024)
CATER: Leveraging LLM to Pioneer a Multidimensional, Reference-Independent Paradigm in Translation Quality Evaluation
di: IIDA, Kurando, et al.
Pubblicazione: (2024)
di: IIDA, Kurando, et al.
Pubblicazione: (2024)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
Do Reasoning Models Enhance Embedding Models?
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
di: Chan, Wun Yu, et al.
Pubblicazione: (2026)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
di: Platt, Nolan, et al.
Pubblicazione: (2025)
di: Platt, Nolan, et al.
Pubblicazione: (2025)
An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT
di: Ma, Chong, et al.
Pubblicazione: (2023)
di: Ma, Chong, et al.
Pubblicazione: (2023)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
di: Khan, Omer Jauhar
Pubblicazione: (2025)
di: Khan, Omer Jauhar
Pubblicazione: (2025)
S2vNTM: Semi-supervised vMF Neural Topic Modeling
di: Xu, Weijie, et al.
Pubblicazione: (2023)
di: Xu, Weijie, et al.
Pubblicazione: (2023)
Reducing Hallucinations in Summarization via Reinforcement Learning with Entity Hallucination Index
di: Katwe, Praveenkumar, et al.
Pubblicazione: (2025)
di: Katwe, Praveenkumar, et al.
Pubblicazione: (2025)
KDSTM: Neural Semi-supervised Topic Modeling with Knowledge Distillation
di: Xu, Weijie, et al.
Pubblicazione: (2023)
di: Xu, Weijie, et al.
Pubblicazione: (2023)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
di: Yu, Haeun, et al.
Pubblicazione: (2024)
di: Yu, Haeun, et al.
Pubblicazione: (2024)
Semantic Synergy: Unlocking Policy Insights and Learning Pathways Through Advanced Skill Mapping
di: Koundouri, Phoebe, et al.
Pubblicazione: (2025)
di: Koundouri, Phoebe, et al.
Pubblicazione: (2025)
Towards Ontology-Enhanced Representation Learning for Large Language Models
di: Ronzano, Francesco, et al.
Pubblicazione: (2024)
di: Ronzano, Francesco, et al.
Pubblicazione: (2024)
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
di: Anderson, Bryce, et al.
Pubblicazione: (2025)
di: Anderson, Bryce, et al.
Pubblicazione: (2025)
The CLEF-2025 CheckThat! Lab: Subjectivity, Fact-Checking, Claim Normalization, and Retrieval
di: Alam, Firoj, et al.
Pubblicazione: (2025)
di: Alam, Firoj, et al.
Pubblicazione: (2025)
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
di: Görge, Rebekka, et al.
Pubblicazione: (2025)
di: Görge, Rebekka, et al.
Pubblicazione: (2025)
ProSwitch: Knowledge-Guided Instruction Tuning to Switch Between Professional and Non-Professional Responses
di: Zong, Chang, et al.
Pubblicazione: (2024)
di: Zong, Chang, et al.
Pubblicazione: (2024)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
di: Marjanović, Sara Vera, et al.
Pubblicazione: (2024)
Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
di: Velampalli, Sirisha, et al.
Pubblicazione: (2025)
di: Velampalli, Sirisha, et al.
Pubblicazione: (2025)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
di: Garg, Saloni, et al.
Pubblicazione: (2026)
di: Garg, Saloni, et al.
Pubblicazione: (2026)
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
di: Juzek, Tom S., et al.
Pubblicazione: (2025)
di: Juzek, Tom S., et al.
Pubblicazione: (2025)
Documenti analoghi
-
Pun Unintended: LLMs and the Illusion of Humor Understanding
di: Zangari, Alessandro, et al.
Pubblicazione: (2025) -
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
di: Zong, Chang, et al.
Pubblicazione: (2024) -
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
di: Wang, Yongjie, et al.
Pubblicazione: (2025) -
The Impact of Role Design in In-Context Learning for Large Language Models
di: Rouzegar, Hamidreza, et al.
Pubblicazione: (2025) -
CLMN: Concept based Language Models via Neural Symbolic Reasoning
di: Yang, Yibo
Pubblicazione: (2025)