Subliminal Signals in Preference Labels
Fuente:
arXiv
Saved in:
| Main Authors: | Magistrali, Isotta, Berdoz, Frédéric, Dauncey, Sam, Wattenhofer, Roger |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Can AI Agents Agree?
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
by: Lorenc, Aleksander, et al.
Published: (2026)
by: Lorenc, Aleksander, et al.
Published: (2026)
Alignment-Aware Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
Text-to-Scene with Large Reasoning Models
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
by: Berdoz, Frédéric, et al.
Published: (2024)
by: Berdoz, Frédéric, et al.
Published: (2024)
Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
by: Fraij, Ahmad, et al.
Published: (2025)
by: Fraij, Ahmad, et al.
Published: (2025)
Approximations to the Fisher Information Metric of Deep Generative Models for Out-Of-Distribution Detection
by: Dauncey, Sam, et al.
Published: (2024)
by: Dauncey, Sam, et al.
Published: (2024)
Benchmarking Music Generation Models and Metrics via Human Preference Studies
by: Grötschla, Florian, et al.
Published: (2025)
by: Grötschla, Florian, et al.
Published: (2025)
Recommender Systems for Democracy: Toward Adversarial Robustness in Voting Advice Applications
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Subliminal Corruption: Mechanisms, Thresholds, and Interpretability
by: Vir, Reya, et al.
Published: (2025)
by: Vir, Reya, et al.
Published: (2025)
On the Expressive Power of GNNs for Boolean Satisfiability
by: Peltonen, Saku, et al.
Published: (2026)
by: Peltonen, Saku, et al.
Published: (2026)
Tiny Transformers Excel at Sentence Compression
by: Belcak, Peter, et al.
Published: (2024)
by: Belcak, Peter, et al.
Published: (2024)
Subliminal Learning is a LoRA Artifact
by: Nief, Todd, et al.
Published: (2026)
by: Nief, Todd, et al.
Published: (2026)
Clone-Robust Weights in Metric Spaces: Handling Redundancy Bias for Benchmark Aggregation
by: Berriaud, Damien, et al.
Published: (2025)
by: Berriaud, Damien, et al.
Published: (2025)
Benchmarking Positional Encodings for GNNs and Graph Transformers
by: Grötschla, Florian, et al.
Published: (2024)
by: Grötschla, Florian, et al.
Published: (2024)
Next Level Message-Passing with Hierarchical Support Graphs
by: Vonessen, Carlos, et al.
Published: (2024)
by: Vonessen, Carlos, et al.
Published: (2024)
Conditional Hallucinations for Image Compression
by: Aczel, Till, et al.
Published: (2024)
by: Aczel, Till, et al.
Published: (2024)
From Message-Passing to Linearized Graph Sequence Models
by: Mathys, Joël, et al.
Published: (2026)
by: Mathys, Joël, et al.
Published: (2026)
Recurrent Deep Differentiable Logic Gate Networks
by: Bührer, Simon, et al.
Published: (2025)
by: Bührer, Simon, et al.
Published: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Efficient Bayesian Inference from Noisy Pairwise Comparisons
by: Aczel, Till, et al.
Published: (2025)
by: Aczel, Till, et al.
Published: (2025)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
by: Askin, Baris, et al.
Published: (2026)
by: Askin, Baris, et al.
Published: (2026)
Light Differentiable Logic Gate Networks
by: Rüttgers, Lukas, et al.
Published: (2025)
by: Rüttgers, Lukas, et al.
Published: (2025)
Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
by: Yousefi, Shakir, et al.
Published: (2025)
by: Yousefi, Shakir, et al.
Published: (2025)
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
by: Gisler, Isaia, et al.
Published: (2026)
by: Gisler, Isaia, et al.
Published: (2026)
Flood and Echo Net: Algorithmically Aligned GNNs that Generalize
by: Mathys, Joël, et al.
Published: (2023)
by: Mathys, Joël, et al.
Published: (2023)
Inductive Transfer Learning for Graph-Based Recommenders
by: Grötschla, Florian, et al.
Published: (2025)
by: Grötschla, Florian, et al.
Published: (2025)
Benchmarking GNNs Using Lightning Network Data
by: Feichtinger, Rainer, et al.
Published: (2024)
by: Feichtinger, Rainer, et al.
Published: (2024)
CoRe-GD: A Hierarchical Framework for Scalable Graph Visualization with GNNs
by: Grötschla, Florian, et al.
Published: (2024)
by: Grötschla, Florian, et al.
Published: (2024)
Assessing Adversarial Robustness of Large Language Models: An Empirical Study
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
by: Cloud, Alex, et al.
Published: (2025)
by: Cloud, Alex, et al.
Published: (2025)
FLIP Reasoning Challenge
by: Plesner, Andreas, et al.
Published: (2025)
by: Plesner, Andreas, et al.
Published: (2025)
Benchmarking Diarization Models
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
Bias beyond Borders: Global Inequalities in AI-Generated Music
by: Solak, Ahmet, et al.
Published: (2025)
by: Solak, Ahmet, et al.
Published: (2025)
High-Fidelity Music Vocoder using Neural Audio Codecs
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
by: Dellali, Amir, et al.
Published: (2025)
by: Dellali, Amir, et al.
Published: (2025)
Similar Items
-
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025) -
Can AI Agents Agree?
by: Berdoz, Frédéric, et al.
Published: (2026) -
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026) -
N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
by: Lorenc, Aleksander, et al.
Published: (2026) -
Alignment-Aware Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)