Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Geng, Yilin, Li, Haonan, Mu, Honglin, Han, Xudong, Baldwin, Timothy, Abend, Omri, Hovy, Eduard, Frermann, Lea |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Measuring Pragmatic Influence in Large Language Model Instructions
por: Geng, Yilin, et al.
Publicado: (2026)
por: Geng, Yilin, et al.
Publicado: (2026)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
por: Wang, Renxi, et al.
Publicado: (2023)
por: Wang, Renxi, et al.
Publicado: (2023)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
por: Berger, Uri, et al.
Publicado: (2025)
por: Berger, Uri, et al.
Publicado: (2025)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
por: Berger, Uri, et al.
Publicado: (2024)
por: Berger, Uri, et al.
Publicado: (2024)
Retain or Reframe? A Computational Framework for the Analysis of Framing in News Articles and Reader Comments
por: Guida, Matteo, et al.
Publicado: (2025)
por: Guida, Matteo, et al.
Publicado: (2025)
LLMs for Argument Mining: Detection, Extraction, and Relationship Classification of pre-defined Arguments in Online Comments
por: Guida, Matteo, et al.
Publicado: (2025)
por: Guida, Matteo, et al.
Publicado: (2025)
Article and Comment Frames Shape the Quality of Online Comments
por: Guida, Matteo, et al.
Publicado: (2026)
por: Guida, Matteo, et al.
Publicado: (2026)
Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs
por: Wagner, Eitan, et al.
Publicado: (2025)
por: Wagner, Eitan, et al.
Publicado: (2025)
Benchmarking Gender and Political Bias in Large Language Models
por: Yang, Jinrui, et al.
Publicado: (2025)
por: Yang, Jinrui, et al.
Publicado: (2025)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
por: Wagner, Eitan, et al.
Publicado: (2024)
por: Wagner, Eitan, et al.
Publicado: (2024)
Connecting the Dots in News Analysis: Bridging the Cross-Disciplinary Disparities in Media Bias and Framing
por: Vallejo, Gisela, et al.
Publicado: (2023)
por: Vallejo, Gisela, et al.
Publicado: (2023)
CONTESTS: a Framework for Consistency Testing of Span Probabilities in Language Models
por: Wagner, Eitan, et al.
Publicado: (2024)
por: Wagner, Eitan, et al.
Publicado: (2024)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
por: Wang, Renxi, et al.
Publicado: (2024)
por: Wang, Renxi, et al.
Publicado: (2024)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
por: Puerto, Haritz, et al.
Publicado: (2026)
por: Puerto, Haritz, et al.
Publicado: (2026)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
por: Elmakies, Avishai, et al.
Publicado: (2025)
por: Elmakies, Avishai, et al.
Publicado: (2025)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
por: Yehudai, Asaf, et al.
Publicado: (2024)
por: Yehudai, Asaf, et al.
Publicado: (2024)
A Language-agnostic Model of Child Language Acquisition
por: Mahon, Louis, et al.
Publicado: (2024)
por: Mahon, Louis, et al.
Publicado: (2024)
Identifying Narrative Patterns and Outliers in Holocaust Testimonies Using Topic Modeling
por: Ifergan, Maxim, et al.
Publicado: (2024)
por: Ifergan, Maxim, et al.
Publicado: (2024)
A Sentiment Consolidation Framework for Meta-Review Generation
por: Li, Miao, et al.
Publicado: (2024)
por: Li, Miao, et al.
Publicado: (2024)
Psychometric Predictive Power of Large Language Models
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
por: Wang, Renxi, et al.
Publicado: (2025)
por: Wang, Renxi, et al.
Publicado: (2025)
Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
por: Mu, Honglin, et al.
Publicado: (2024)
por: Mu, Honglin, et al.
Publicado: (2024)
Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning
por: Wagner, Eitan, et al.
Publicado: (2024)
por: Wagner, Eitan, et al.
Publicado: (2024)
Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models
por: Apartsin, Alexander, et al.
Publicado: (2026)
por: Apartsin, Alexander, et al.
Publicado: (2026)
$T^5Score$: A Methodology for Automatically Assessing the Quality of LLM Generated Multi-Document Topic Sets
por: Trainin, Itamar, et al.
Publicado: (2024)
por: Trainin, Itamar, et al.
Publicado: (2024)
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
por: Mu, Honglin, et al.
Publicado: (2025)
por: Mu, Honglin, et al.
Publicado: (2025)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
por: Ifergan, Maxim, et al.
Publicado: (2024)
por: Ifergan, Maxim, et al.
Publicado: (2024)
Assessing the Role of Lexical Semantics in Cross-lingual Transfer through Controlled Manipulations
por: Ilani, Roy, et al.
Publicado: (2024)
por: Ilani, Roy, et al.
Publicado: (2024)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
por: Curran, Damian, et al.
Publicado: (2025)
por: Curran, Damian, et al.
Publicado: (2025)
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
por: De Michele, William, et al.
Publicado: (2025)
por: De Michele, William, et al.
Publicado: (2025)
Factuality Challenges in the Era of Large Language Models
por: Augenstein, Isabelle, et al.
Publicado: (2023)
por: Augenstein, Isabelle, et al.
Publicado: (2023)
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
por: Wang, Yuxia, et al.
Publicado: (2024)
por: Wang, Yuxia, et al.
Publicado: (2024)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
por: Russo, Giuseppe, et al.
Publicado: (2025)
por: Russo, Giuseppe, et al.
Publicado: (2025)
Narrative Media Framing in Political Discourse
por: Otmakhova, Yulia, et al.
Publicado: (2025)
por: Otmakhova, Yulia, et al.
Publicado: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
por: Röttger, Paul, et al.
Publicado: (2024)
por: Röttger, Paul, et al.
Publicado: (2024)
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
por: Wang, Shuhe, et al.
Publicado: (2024)
por: Wang, Shuhe, et al.
Publicado: (2024)
Large Language Model Reasoning Failures
por: Song, Peiyang, et al.
Publicado: (2026)
por: Song, Peiyang, et al.
Publicado: (2026)
Generating Benchmarks for Factuality Evaluation of Language Models
por: Muhlgay, Dor, et al.
Publicado: (2023)
por: Muhlgay, Dor, et al.
Publicado: (2023)
Are UFOs Driving Innovation? The Illusion of Causality in Large Language Models
por: Carro, María Victoria, et al.
Publicado: (2024)
por: Carro, María Victoria, et al.
Publicado: (2024)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
por: Don-Yehiya, Shachar, et al.
Publicado: (2024)
por: Don-Yehiya, Shachar, et al.
Publicado: (2024)
Ejemplares similares
-
Measuring Pragmatic Influence in Large Language Model Instructions
por: Geng, Yilin, et al.
Publicado: (2026) -
Demystifying Instruction Mixing for Fine-tuning Large Language Models
por: Wang, Renxi, et al.
Publicado: (2023) -
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
por: Berger, Uri, et al.
Publicado: (2025) -
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
por: Berger, Uri, et al.
Publicado: (2024) -
Retain or Reframe? A Computational Framework for the Analysis of Framing in News Articles and Reader Comments
por: Guida, Matteo, et al.
Publicado: (2025)