Subliminal Signals in Preference Labels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Magistrali, Isotta, Berdoz, Frédéric, Dauncey, Sam, Wattenhofer, Roger |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Steering Pretrained Drafters during Speculative Decoding
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
Can AI Agents Agree?
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
Reasoning Boosts Opinion Alignment in LLMs
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
von: Lorenc, Aleksander, et al.
Veröffentlicht: (2026)
von: Lorenc, Aleksander, et al.
Veröffentlicht: (2026)
Alignment-Aware Decoding
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
WorldSpeech: A Multilingual Speech Corpus from Around the World
von: Asonitis, Antonis, et al.
Veröffentlicht: (2026)
von: Asonitis, Antonis, et al.
Veröffentlicht: (2026)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Text-to-Scene with Large Reasoning Models
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2024)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2024)
Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
von: Fraij, Ahmad, et al.
Veröffentlicht: (2025)
von: Fraij, Ahmad, et al.
Veröffentlicht: (2025)
Approximations to the Fisher Information Metric of Deep Generative Models for Out-Of-Distribution Detection
von: Dauncey, Sam, et al.
Veröffentlicht: (2024)
von: Dauncey, Sam, et al.
Veröffentlicht: (2024)
Benchmarking Music Generation Models and Metrics via Human Preference Studies
von: Grötschla, Florian, et al.
Veröffentlicht: (2025)
von: Grötschla, Florian, et al.
Veröffentlicht: (2025)
Recommender Systems for Democracy: Toward Adversarial Robustness in Voting Advice Applications
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
Subliminal Corruption: Mechanisms, Thresholds, and Interpretability
von: Vir, Reya, et al.
Veröffentlicht: (2025)
von: Vir, Reya, et al.
Veröffentlicht: (2025)
On the Expressive Power of GNNs for Boolean Satisfiability
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
Tiny Transformers Excel at Sentence Compression
von: Belcak, Peter, et al.
Veröffentlicht: (2024)
von: Belcak, Peter, et al.
Veröffentlicht: (2024)
Subliminal Learning is a LoRA Artifact
von: Nief, Todd, et al.
Veröffentlicht: (2026)
von: Nief, Todd, et al.
Veröffentlicht: (2026)
Clone-Robust Weights in Metric Spaces: Handling Redundancy Bias for Benchmark Aggregation
von: Berriaud, Damien, et al.
Veröffentlicht: (2025)
von: Berriaud, Damien, et al.
Veröffentlicht: (2025)
Benchmarking Positional Encodings for GNNs and Graph Transformers
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
Next Level Message-Passing with Hierarchical Support Graphs
von: Vonessen, Carlos, et al.
Veröffentlicht: (2024)
von: Vonessen, Carlos, et al.
Veröffentlicht: (2024)
Conditional Hallucinations for Image Compression
von: Aczel, Till, et al.
Veröffentlicht: (2024)
von: Aczel, Till, et al.
Veröffentlicht: (2024)
From Message-Passing to Linearized Graph Sequence Models
von: Mathys, Joël, et al.
Veröffentlicht: (2026)
von: Mathys, Joël, et al.
Veröffentlicht: (2026)
Recurrent Deep Differentiable Logic Gate Networks
von: Bührer, Simon, et al.
Veröffentlicht: (2025)
von: Bührer, Simon, et al.
Veröffentlicht: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
Efficient Bayesian Inference from Noisy Pairwise Comparisons
von: Aczel, Till, et al.
Veröffentlicht: (2025)
von: Aczel, Till, et al.
Veröffentlicht: (2025)
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
von: Askin, Baris, et al.
Veröffentlicht: (2026)
von: Askin, Baris, et al.
Veröffentlicht: (2026)
Light Differentiable Logic Gate Networks
von: Rüttgers, Lukas, et al.
Veröffentlicht: (2025)
von: Rüttgers, Lukas, et al.
Veröffentlicht: (2025)
Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
von: Yousefi, Shakir, et al.
Veröffentlicht: (2025)
von: Yousefi, Shakir, et al.
Veröffentlicht: (2025)
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
von: Gisler, Isaia, et al.
Veröffentlicht: (2026)
von: Gisler, Isaia, et al.
Veröffentlicht: (2026)
Flood and Echo Net: Algorithmically Aligned GNNs that Generalize
von: Mathys, Joël, et al.
Veröffentlicht: (2023)
von: Mathys, Joël, et al.
Veröffentlicht: (2023)
Inductive Transfer Learning for Graph-Based Recommenders
von: Grötschla, Florian, et al.
Veröffentlicht: (2025)
von: Grötschla, Florian, et al.
Veröffentlicht: (2025)
Benchmarking GNNs Using Lightning Network Data
von: Feichtinger, Rainer, et al.
Veröffentlicht: (2024)
von: Feichtinger, Rainer, et al.
Veröffentlicht: (2024)
CoRe-GD: A Hierarchical Framework for Scalable Graph Visualization with GNNs
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
Assessing Adversarial Robustness of Large Language Models: An Empirical Study
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
von: Cloud, Alex, et al.
Veröffentlicht: (2025)
von: Cloud, Alex, et al.
Veröffentlicht: (2025)
FLIP Reasoning Challenge
von: Plesner, Andreas, et al.
Veröffentlicht: (2025)
von: Plesner, Andreas, et al.
Veröffentlicht: (2025)
Benchmarking Diarization Models
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Bias beyond Borders: Global Inequalities in AI-Generated Music
von: Solak, Ahmet, et al.
Veröffentlicht: (2025)
von: Solak, Ahmet, et al.
Veröffentlicht: (2025)
High-Fidelity Music Vocoder using Neural Audio Codecs
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
von: Dellali, Amir, et al.
Veröffentlicht: (2025)
von: Dellali, Amir, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Steering Pretrained Drafters during Speculative Decoding
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025) -
Can AI Agents Agree?
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026) -
Reasoning Boosts Opinion Alignment in LLMs
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026) -
N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
von: Lorenc, Aleksander, et al.
Veröffentlicht: (2026) -
Alignment-Aware Decoding
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)