Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
Fuente:
arXiv
Guardado en:
| Autores principales: | Aden-Ali, Ishaq, Golowich, Noah, Liu, Allen, Shetty, Abhishek, Moitra, Ankur, Haghtalab, Nika |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Sequences of Logits Reveal the Low Rank Structure of Language Models
por: Golowich, Noah, et al.
Publicado: (2025)
por: Golowich, Noah, et al.
Publicado: (2025)
Edit Distance Robust Watermarks via Indexing Pseudorandom Codes
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Provably Learning from Modern Language Models via Low Logit Rank
por: Golowich, Noah, et al.
Publicado: (2025)
por: Golowich, Noah, et al.
Publicado: (2025)
Smooth Nash Equilibria: Algorithms and Complexity
por: Daskalakis, Constantinos, et al.
Publicado: (2023)
por: Daskalakis, Constantinos, et al.
Publicado: (2023)
Model Stealing for Any Low-Rank Language Model
por: Liu, Allen, et al.
Publicado: (2024)
por: Liu, Allen, et al.
Publicado: (2024)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
por: Askin, Baris, et al.
Publicado: (2026)
por: Askin, Baris, et al.
Publicado: (2026)
The Role of Sparsity for Length Generalization in Transformers
por: Golowich, Noah, et al.
Publicado: (2025)
por: Golowich, Noah, et al.
Publicado: (2025)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
por: Haussmann, Aden
Publicado: (2025)
por: Haussmann, Aden
Publicado: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
por: Halawi, Danny, et al.
Publicado: (2024)
por: Halawi, Danny, et al.
Publicado: (2024)
Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
On Learning Parities with Dependent Noise
por: Golowich, Noah, et al.
Publicado: (2024)
por: Golowich, Noah, et al.
Publicado: (2024)
Is Efficient PAC Learning Possible with an Oracle That Responds 'Yes' or 'No'?
por: Daskalakis, Constantinos, et al.
Publicado: (2024)
por: Daskalakis, Constantinos, et al.
Publicado: (2024)
On Linear Representations and Pretraining Data Frequency in Language Models
por: Merullo, Jack, et al.
Publicado: (2025)
por: Merullo, Jack, et al.
Publicado: (2025)
Your Transformer is Secretly Linear
por: Razzhigaev, Anton, et al.
Publicado: (2024)
por: Razzhigaev, Anton, et al.
Publicado: (2024)
From Style to Facts: Mapping the Boundaries of Knowledge Injection with Finetuning
por: Zhao, Eric, et al.
Publicado: (2025)
por: Zhao, Eric, et al.
Publicado: (2025)
HALT: Hallucination Assessment via Log-probs as Time series
por: Shapiro, Ahmad, et al.
Publicado: (2026)
por: Shapiro, Ahmad, et al.
Publicado: (2026)
Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference
por: Golowich, Noah, et al.
Publicado: (2026)
por: Golowich, Noah, et al.
Publicado: (2026)
Valid Text-to-SQL Generation with Unification-based DeepStochLog
por: Jiao, Ying, et al.
Publicado: (2025)
por: Jiao, Ying, et al.
Publicado: (2025)
SelfElicit: Your Language Model Secretly Knows Where is the Relevant Evidence
por: Liu, Zhining, et al.
Publicado: (2025)
por: Liu, Zhining, et al.
Publicado: (2025)
Subliminal Learning Is Steering Vector Distillation
por: Blank, Camila, et al.
Publicado: (2026)
por: Blank, Camila, et al.
Publicado: (2026)
Improve Large Language Model Systems with User Logs
por: Wang, Changyue, et al.
Publicado: (2026)
por: Wang, Changyue, et al.
Publicado: (2026)
Prove Your Point!: Bringing Proof-Enhancement Principles to Argumentative Essay Generation
por: Xiao, Ruiyu, et al.
Publicado: (2024)
por: Xiao, Ruiyu, et al.
Publicado: (2024)
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
por: Rathore, Darshita, et al.
Publicado: (2025)
por: Rathore, Darshita, et al.
Publicado: (2025)
The Coverage Principle: How Pre-Training Enables Post-Training
por: Chen, Fan, et al.
Publicado: (2025)
por: Chen, Fan, et al.
Publicado: (2025)
Chatting with Logs: An exploratory study on Finetuning LLMs for LogQL
por: Seshagiri, Vishwanath, et al.
Publicado: (2024)
por: Seshagiri, Vishwanath, et al.
Publicado: (2024)
sDPO: Don't Use Your Data All at Once
por: Kim, Dahyun, et al.
Publicado: (2024)
por: Kim, Dahyun, et al.
Publicado: (2024)
Pastiche Novel Generation Creating: Fan Fiction You Love in Your Favorite Author's Style
por: Han, Xueran, et al.
Publicado: (2025)
por: Han, Xueran, et al.
Publicado: (2025)
Are Your LLMs Capable of Stable Reasoning?
por: Liu, Junnan, et al.
Publicado: (2024)
por: Liu, Junnan, et al.
Publicado: (2024)
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
por: Zhussip, Magauiya, et al.
Publicado: (2025)
por: Zhussip, Magauiya, et al.
Publicado: (2025)
The Hidden Game Problem
por: Buzaglo, Gon, et al.
Publicado: (2025)
por: Buzaglo, Gon, et al.
Publicado: (2025)
Syntriever: How to Train Your Retriever with Synthetic Data from LLMs
por: Kim, Minsang, et al.
Publicado: (2025)
por: Kim, Minsang, et al.
Publicado: (2025)
Is Your LLM Outdated? A Deep Look at Temporal Generalization
por: Zhu, Chenghao, et al.
Publicado: (2024)
por: Zhu, Chenghao, et al.
Publicado: (2024)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
por: Murphy, Alexander, et al.
Publicado: (2025)
por: Murphy, Alexander, et al.
Publicado: (2025)
Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text
por: Kempton, Tom, et al.
Publicado: (2026)
por: Kempton, Tom, et al.
Publicado: (2026)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
por: Sankaranarayanan, Aruna, et al.
Publicado: (2025)
por: Sankaranarayanan, Aruna, et al.
Publicado: (2025)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
por: Liu, Xianyang, et al.
Publicado: (2025)
por: Liu, Xianyang, et al.
Publicado: (2025)
Not All Synthetic Data Is Yours to Learn From
por: Alemohammad, Sina, et al.
Publicado: (2026)
por: Alemohammad, Sina, et al.
Publicado: (2026)
How to Engage Your Readers? Generating Guiding Questions to Promote Active Reading
por: Cui, Peng, et al.
Publicado: (2024)
por: Cui, Peng, et al.
Publicado: (2024)
Ejemplares similares
-
Linear Bellman Completeness Suffices for Efficient Online Reinforcement Learning with Few Actions
por: Golowich, Noah, et al.
Publicado: (2024) -
The Role of Inherent Bellman Error in Offline Reinforcement Learning with Linear Function Approximation
por: Golowich, Noah, et al.
Publicado: (2024) -
Sequences of Logits Reveal the Low Rank Structure of Language Models
por: Golowich, Noah, et al.
Publicado: (2025) -
Edit Distance Robust Watermarks via Indexing Pseudorandom Codes
por: Golowich, Noah, et al.
Publicado: (2024) -
Provably Learning from Modern Language Models via Low Logit Rank
por: Golowich, Noah, et al.
Publicado: (2025)