Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Fuente:
arXiv
Guardado en:
| Autores principales: | Tan, Daniel, Woodruff, Anders, Warncke, Niels, Jose, Arun, Riché, Maxime, Africa, David Demitri, Taylor, Mia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025)
por: Africa, David Demitri
Publicado: (2025)
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
por: Wichers, Nevan, et al.
Publicado: (2025)
por: Wichers, Nevan, et al.
Publicado: (2025)
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
por: Ivanov, Igor, et al.
Publicado: (2026)
por: Ivanov, Igor, et al.
Publicado: (2026)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
por: Betley, Jan, et al.
Publicado: (2025)
por: Betley, Jan, et al.
Publicado: (2025)
Steering Awareness: Detecting Activation Steering from Within
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)
Does Self-Evaluation Enable Wireheading in Language Models?
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Consistency Training while Mitigating Obfuscation via Rate Matching
por: Imran, Sohaib, et al.
Publicado: (2026)
por: Imran, Sohaib, et al.
Publicado: (2026)
Learning Dynamics of Meta-Learning in Small Model Pretraining
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
por: Weiss, Yuval, et al.
Publicado: (2025)
por: Weiss, Yuval, et al.
Publicado: (2025)
Zwei neue Sandbienen aus der Ukraine und aus Ungarn (Hym. Apoidea)
por: Warncke, Klaus
Publicado: (1972)
por: Warncke, Klaus
Publicado: (1972)
Plan for the Development of Library Service in Montana.
por: Warncke, Ruth
Publicado: (1965)
por: Warncke, Ruth
Publicado: (1965)
Analyzing Your Community: Basis for Building Library Service.
por: Warncke, Ruth
Publicado: (1974)
por: Warncke, Ruth
Publicado: (1974)
Planning Library Workshops and Institutes.
por: Warncke, Ruth
Publicado: (1976)
por: Warncke, Ruth
Publicado: (1976)
Learning Modular Exponentiation with Transformers
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
por: Africa, David Demitri, et al.
Publicado: (2025)
por: Africa, David Demitri, et al.
Publicado: (2025)
ALA Report to IFLA for 1970-1971
por: Clift, David H., et al.
Publicado: (1972)
por: Clift, David H., et al.
Publicado: (1972)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
por: Martinez, Richard Diehl, et al.
Publicado: (2025)
por: Martinez, Richard Diehl, et al.
Publicado: (2025)
LLMs cannot find reasoning errors, but can correct them given the error location
por: Tyen, Gladys, et al.
Publicado: (2023)
por: Tyen, Gladys, et al.
Publicado: (2023)
Mixed modular perverse sheaves on affine flag varieties and Koszul duality
por: Riche, Simon
Publicado: (2024)
por: Riche, Simon
Publicado: (2024)
Emancipation's Daughters
por: Richardson, Riché
Publicado: (2021)
por: Richardson, Riché
Publicado: (2021)
Some applications of the geometric Satake equivalence to modular representation theory
por: Riche, Simon
Publicado: (2024)
por: Riche, Simon
Publicado: (2024)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
por: Cencerrado, Iván Vicente Moreno, et al.
Publicado: (2025)
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
por: Puerto, Haritz, et al.
Publicado: (2024)
por: Puerto, Haritz, et al.
Publicado: (2024)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
por: Naderi, Nariman, et al.
Publicado: (2025)
por: Naderi, Nariman, et al.
Publicado: (2025)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
por: Agarwal, Utkarsh, et al.
Publicado: (2024)
por: Agarwal, Utkarsh, et al.
Publicado: (2024)
Terpenes and Terpenoids: How can we use them?
por: Jay Hanssens, et al.
Publicado: (2025)
por: Jay Hanssens, et al.
Publicado: (2025)
Useful Public Spending, Taylor Principle, and Macroeconomic Instability
por: Antoine Le Riche
Publicado: (2025)
por: Antoine Le Riche
Publicado: (2025)
LLMs can see and hear without any training
por: Ashutosh, Kumar, et al.
Publicado: (2025)
por: Ashutosh, Kumar, et al.
Publicado: (2025)
MALicious INTent Dataset and Inoculating LLMs for Enhanced Disinformation Detection
por: Modzelewski, Arkadiusz, et al.
Publicado: (2026)
por: Modzelewski, Arkadiusz, et al.
Publicado: (2026)
Implementing surrogate goals for safer bargaining in LLM-based agents
por: Oesterheld, Caspar, et al.
Publicado: (2026)
por: Oesterheld, Caspar, et al.
Publicado: (2026)
Analogies and differences between the logic behind statistical hypothesis testing and proofs by contradiction: What can we learn from them?
por: Maria Cristina Amoretti, et al.
Publicado: (2026)
por: Maria Cristina Amoretti, et al.
Publicado: (2026)
Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
por: Zhang, Tanghaoran, et al.
Publicado: (2024)
por: Zhang, Tanghaoran, et al.
Publicado: (2024)
Inoculation and Co‐Inoculation With Plant Growth–Promoting Bacteria in Chickpea: Physiological Aspects and Plant Growth
por: Karla Sabrina Magalhães Andrade Padilha, et al.
Publicado: (2026)
por: Karla Sabrina Magalhães Andrade Padilha, et al.
Publicado: (2026)
Elicitive Curricular Development
por: Echavarría Alvarez, Josefina, et al.
Publicado: (2020)
por: Echavarría Alvarez, Josefina, et al.
Publicado: (2020)
On two modular geometric realizations of an affine Hecke algebra
por: Bezrukavnikov, Roman, et al.
Publicado: (2024)
por: Bezrukavnikov, Roman, et al.
Publicado: (2024)
Equivariant Koszul Duality, Modular Category $\mathcal{O}$, and Periodic Kazhdan--Lusztig Polynomials
por: Riche, Simon, et al.
Publicado: (2025)
por: Riche, Simon, et al.
Publicado: (2025)
On multi-graded Proj schemes
por: Mayeux, Arnaud, et al.
Publicado: (2023)
por: Mayeux, Arnaud, et al.
Publicado: (2023)
Modular affine Hecke category and regular centralizer
por: Bezrukavnikov, Roman, et al.
Publicado: (2022)
por: Bezrukavnikov, Roman, et al.
Publicado: (2022)
Koszul duality for Coxeter groups
por: Riche, Simon, et al.
Publicado: (2023)
por: Riche, Simon, et al.
Publicado: (2023)
Conversational Inoculation to Enhance Resistance to Misinformation
por: Szabó, Dániel, et al.
Publicado: (2026)
por: Szabó, Dániel, et al.
Publicado: (2026)
Ejemplares similares
-
Identifying a Circuit for Verb Conjugation in GPT-2
por: Africa, David Demitri
Publicado: (2025) -
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
por: Wichers, Nevan, et al.
Publicado: (2025) -
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
por: Ivanov, Igor, et al.
Publicado: (2026) -
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
por: Betley, Jan, et al.
Publicado: (2025) -
Steering Awareness: Detecting Activation Steering from Within
por: Rivera, Joshua Fonseca, et al.
Publicado: (2025)