Gespeichert in:
| Hauptverfasser: | Treutlein, Johannes, Choi, Dami, Betley, Jan, Marks, Samuel, Anil, Cem, Grosse, Roger, Evans, Owain |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2406.14546 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
von: Taylor, Mia, et al.
Veröffentlicht: (2025)
von: Taylor, Mia, et al.
Veröffentlicht: (2025)
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
von: Chua, James, et al.
Veröffentlicht: (2026)
von: Chua, James, et al.
Veröffentlicht: (2026)
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
von: Dubiński, Jan, et al.
Veröffentlicht: (2026)
von: Dubiński, Jan, et al.
Veröffentlicht: (2026)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
von: Cloud, Alex, et al.
Veröffentlicht: (2025)
von: Cloud, Alex, et al.
Veröffentlicht: (2025)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025)
von: Betley, Jan, et al.
Veröffentlicht: (2025)
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
von: Chua, James, et al.
Veröffentlicht: (2025)
von: Chua, James, et al.
Veröffentlicht: (2025)
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
von: Laine, Rudolf, et al.
Veröffentlicht: (2024)
Lessons from Studying Two-Hop Latent Reasoning
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
von: Balesni, Mikita, et al.
Veröffentlicht: (2024)
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
von: Karvonen, Adam, et al.
Veröffentlicht: (2025)
Are DeepSeek R1 And Other Reasoning Models More Faithful?
von: Chua, James, et al.
Veröffentlicht: (2025)
von: Chua, James, et al.
Veröffentlicht: (2025)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
Can Language Models Explain Their Own Classification Behavior?
von: Sherburn, Dane, et al.
Veröffentlicht: (2024)
von: Sherburn, Dane, et al.
Veröffentlicht: (2024)
Connecting the Dots: Inferring Patent Phrase Similarity with Retrieved Phrase Graphs
von: Peng, Zhuoyi, et al.
Veröffentlicht: (2024)
von: Peng, Zhuoyi, et al.
Veröffentlicht: (2024)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
Inferring the presence and abundance of rare waterbirds species from scarce data
von: Bricout, Barbara, et al.
Veröffentlicht: (2026)
von: Bricout, Barbara, et al.
Veröffentlicht: (2026)
Can We Infer Confidential Properties of Training Data from LLMs?
von: Huang, Pengrun, et al.
Veröffentlicht: (2025)
von: Huang, Pengrun, et al.
Veröffentlicht: (2025)
Cylinder decompositions on geometric armadillo tails
von: Lee, Dami, et al.
Veröffentlicht: (2024)
von: Lee, Dami, et al.
Veröffentlicht: (2024)
Estrutura de propriedade no Brasil: Evidências empíricas no grau de concentração acionária
von: Anamélia Borges Tannus Dami
Veröffentlicht: (2023)
von: Anamélia Borges Tannus Dami
Veröffentlicht: (2023)
ESTRUTURA DE PROPRIEDADE NO BRASIL: EVIDÊNCIAS EMPÍRICAS NO GRAU DE CONCENTRAÇÃO ACIONÁRIA
von: Anamélia Borges Tannus Dami
Veröffentlicht: (2007)
von: Anamélia Borges Tannus Dami
Veröffentlicht: (2007)
Training Data Attribution via Approximate Unrolled Differentiation
von: Bae, Juhan, et al.
Veröffentlicht: (2024)
von: Bae, Juhan, et al.
Veröffentlicht: (2024)
Connecting the Dots in News Analysis: Bridging the Cross-Disciplinary Disparities in Media Bias and Framing
von: Vallejo, Gisela, et al.
Veröffentlicht: (2023)
von: Vallejo, Gisela, et al.
Veröffentlicht: (2023)
Discrete Vector Bundles with Connection
von: Berwick-Evans, Daniel, et al.
Veröffentlicht: (2021)
von: Berwick-Evans, Daniel, et al.
Veröffentlicht: (2021)
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
von: Marks, Samuel, et al.
Veröffentlicht: (2023)
von: Marks, Samuel, et al.
Veröffentlicht: (2023)
Effective Faraday interaction between light and nuclear spins of Helium-3 in its ground state: a semiclassical study
von: Fadel, Matteo, et al.
Veröffentlicht: (2024)
von: Fadel, Matteo, et al.
Veröffentlicht: (2024)
Kinetic analysis of phase transformations during continuous heating: Crystallization of glass-forming liquids
von: Houghton, Owain S.
Veröffentlicht: (2025)
von: Houghton, Owain S.
Veröffentlicht: (2025)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
Resonant Structures in $p{}^7\mathrm{Be}$ Scattering and Their Connection to the Astrophysical $S$-Factor
von: Khachi, Anil
Veröffentlicht: (2025)
von: Khachi, Anil
Veröffentlicht: (2025)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
von: Berglund, Lukas, et al.
Veröffentlicht: (2023)
On Verbalized Confidence Scores for LLMs
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
von: Yang, Daniel, et al.
Veröffentlicht: (2024)
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
von: Shilov, Igor, et al.
Veröffentlicht: (2025)
von: Shilov, Igor, et al.
Veröffentlicht: (2025)
Bayesian Tensor Decomposition for Clustering Latent Symptom Profiles for Verbal Autopsy Data
von: Yu Zhu, et al.
Veröffentlicht: (2026)
von: Yu Zhu, et al.
Veröffentlicht: (2026)
Negation Neglect: When models fail to learn negations in training
von: Mayne, Harry, et al.
Veröffentlicht: (2026)
von: Mayne, Harry, et al.
Veröffentlicht: (2026)
The growth of the mussel Mytilus californianus
von: Richards, Owain Westmacott
Veröffentlicht: (1928)
von: Richards, Owain Westmacott
Veröffentlicht: (1928)
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language
von: Requeima, James, et al.
Veröffentlicht: (2024)
von: Requeima, James, et al.
Veröffentlicht: (2024)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
von: Hubinger, Evan, et al.
Veröffentlicht: (2024)
von: Hubinger, Evan, et al.
Veröffentlicht: (2024)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
von: Chen, Runjin, et al.
Veröffentlicht: (2025)
von: Chen, Runjin, et al.
Veröffentlicht: (2025)
Put Aside Your Pencil: How Talk Becomes Writing Through Verbal Rehearsal
von: Kristen I. Evans
Veröffentlicht: (2026)
von: Kristen I. Evans
Veröffentlicht: (2026)
Ähnliche Einträge
-
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
von: Taylor, Mia, et al.
Veröffentlicht: (2025) -
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious
von: Chua, James, et al.
Veröffentlicht: (2026) -
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
von: Dubiński, Jan, et al.
Veröffentlicht: (2026) -
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
von: Betley, Jan, et al.
Veröffentlicht: (2025) -
Tell me about yourself: LLMs are aware of their learned behaviors
von: Betley, Jan, et al.
Veröffentlicht: (2025)