Saved in:
| Main Authors: | Wen, Jiaxin, Ankner, Zachary, Somani, Arushi, Hase, Peter, Marks, Samuel, Goldman-Wetzler, Jacob, Petrini, Linda, Sleight, Henry, Burns, Collin, He, He, Feng, Shi, Perez, Ethan, Leike, Jan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.10139 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025)
by: Gema, Aryo Pradipta, et al.
Published: (2025)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
by: Shilov, Igor, et al.
Published: (2025)
by: Shilov, Igor, et al.
Published: (2025)
Speeding up and reducing memory usage for scientific machine learning via mixed precision
by: Hayford, Joel, et al.
Published: (2024)
by: Hayford, Joel, et al.
Published: (2024)
Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion
by: Parthasarathy, Rishab, et al.
Published: (2024)
by: Parthasarathy, Rishab, et al.
Published: (2024)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
by: Peng, Alwin, et al.
Published: (2024)
by: Peng, Alwin, et al.
Published: (2024)
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
by: Slocum, Stewart, et al.
Published: (2025)
by: Slocum, Stewart, et al.
Published: (2025)
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
by: Cloud, Alex, et al.
Published: (2024)
by: Cloud, Alex, et al.
Published: (2024)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
by: Youstra, Jack, et al.
Published: (2025)
by: Youstra, Jack, et al.
Published: (2025)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
by: Hägele, Alexander, et al.
Published: (2026)
by: Hägele, Alexander, et al.
Published: (2026)
Language Models Learn to Mislead Humans via RLHF
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
Unsupervised decoding of encoded reasoning using language model interpretability
by: Fang, Ching, et al.
Published: (2025)
by: Fang, Ching, et al.
Published: (2025)
Ethos des literarischen Schreibens
by: Hase, Jan
Published: (2024)
by: Hase, Jan
Published: (2024)
Paul Bowman, Mythologies of Martial Arts (Martial Arts Studies, 2), London/New York: Rowman and Littlefield, 2017, xxiii + 186 p.
by: Sixt Wetzler
Published: (2017)
by: Sixt Wetzler
Published: (2017)
Fight Books in Comparative Perspective. An Introduction
by: Sixt Wetzler
Published: (2020)
by: Sixt Wetzler
Published: (2020)
“Your Kung Fu is very good, Master Fiore!” Asian and European fight books in comparison
by: Sixt Wetzler
Published: (2016)
by: Sixt Wetzler
Published: (2016)
Feminist mental health activism in England c. 1968–95. By K.Mahoney, Manchester: Manchester University Press. 2023. pp. 280. £85.00 (cloth), £85.00 (ebk). ISBN: 9781526162267
by: Sara Wetzler
Published: (2024)
by: Sara Wetzler
Published: (2024)
Critique-out-Loud Reward Models
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
Dynamic Masking Rate Schedules for MLM Pretraining
by: Ankner, Zachary, et al.
Published: (2023)
by: Ankner, Zachary, et al.
Published: (2023)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
by: Guo, Shiyuan, et al.
Published: (2025)
by: Guo, Shiyuan, et al.
Published: (2025)
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
by: Ensign, Danielle, et al.
Published: (2025)
by: Ensign, Danielle, et al.
Published: (2025)
No Difference in Pullout Strength Between a Bio‐inductive Implant and a Semitendinosus Tendon Graft in a Biomechanical Study of Medial Patellofemoral Ligament Repair Augmentation
by: Austin Wetzler, et al.
Published: (2024)
by: Austin Wetzler, et al.
Published: (2024)
Eliciting Secret Knowledge from Language Models
by: Cywiński, Bartosz, et al.
Published: (2025)
by: Cywiński, Bartosz, et al.
Published: (2025)
Discovering Latent Knowledge in Language Models Without Supervision
by: Burns, Collin, et al.
Published: (2022)
by: Burns, Collin, et al.
Published: (2022)
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
by: Casademunt, Helena, et al.
Published: (2026)
by: Casademunt, Helena, et al.
Published: (2026)
Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
Limit-Computable Grains of Truth for Arbitrary Computable Extensive-Form (Un)Known Games
by: Wyeth, Cole, et al.
Published: (2025)
by: Wyeth, Cole, et al.
Published: (2025)
Excess Description Length of Learning Generalizable Predictors
by: Donoway, Elizabeth, et al.
Published: (2026)
by: Donoway, Elizabeth, et al.
Published: (2026)
SMIXAE: Towards Unsupervised Manifold Discovery in Language Models
by: Francel, Collin
Published: (2026)
by: Francel, Collin
Published: (2026)
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
by: Wichers, Nevan, et al.
Published: (2025)
by: Wichers, Nevan, et al.
Published: (2025)
The Impact of Acquisition on Product Quality in the Console Gaming Industry
by: Somani, Shivam
Published: (2024)
by: Somani, Shivam
Published: (2024)
Towards Verifiable Transformers: Solver-Checkable Circuit Explanations
by: Somani, Neel
Published: (2026)
by: Somani, Neel
Published: (2026)
Looking Inward: Language Models Can Learn About Themselves by Introspection
by: Binder, Felix J, et al.
Published: (2024)
by: Binder, Felix J, et al.
Published: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
A High-Order Conformal FEM for Multidimensional Nonlinear Collisional Breakage Equations: Analysis and Computation
by: Arushi, Arushi, et al.
Published: (2026)
by: Arushi, Arushi, et al.
Published: (2026)
Forecasting Rare Language Model Behaviors
by: Jones, Erik, et al.
Published: (2025)
by: Jones, Erik, et al.
Published: (2025)
(Non-)Conserved Currents and Cosmological Correlators
by: Sleight, Charlotte, et al.
Published: (2025)
by: Sleight, Charlotte, et al.
Published: (2025)
Celestial Holography Revisited
by: Sleight, Charlotte, et al.
Published: (2023)
by: Sleight, Charlotte, et al.
Published: (2023)
Similar Items
-
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025) -
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025) -
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
by: Shilov, Igor, et al.
Published: (2025) -
Speeding up and reducing memory usage for scientific machine learning via mixed precision
by: Hayford, Joel, et al.
Published: (2024) -
Vid3D: Synthesis of Dynamic 3D Scenes using 2D Video Diffusion
by: Parthasarathy, Rishab, et al.
Published: (2024)