Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lys, Jonathan, Gripon, Vincent, Pasdeloup, Bastien, Marmoret, Axel, Mauch, Lukas, Cardinaux, Fabien, Hacene, Ghouthi Boukli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training
von: Lys, Jonathan, et al.
Veröffentlicht: (2026)
von: Lys, Jonathan, et al.
Veröffentlicht: (2026)
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding
von: Lys, Jonathan, et al.
Veröffentlicht: (2026)
von: Lys, Jonathan, et al.
Veröffentlicht: (2026)
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
von: Bendou, Yassir, et al.
Veröffentlicht: (2024)
von: Bendou, Yassir, et al.
Veröffentlicht: (2024)
A Novel Benchmark for Few-Shot Semantic Segmentation in the Era of Foundation Models
von: Bensaid, Reda, et al.
Veröffentlicht: (2024)
von: Bensaid, Reda, et al.
Veröffentlicht: (2024)
GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
von: Choudhary, Anand, et al.
Veröffentlicht: (2025)
von: Choudhary, Anand, et al.
Veröffentlicht: (2025)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
von: Marmoret, Axel, et al.
Veröffentlicht: (2025)
von: Marmoret, Axel, et al.
Veröffentlicht: (2025)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
von: Marmoret, Axel
Veröffentlicht: (2026)
von: Marmoret, Axel
Veröffentlicht: (2026)
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
von: Kögel, Fabian, et al.
Veröffentlicht: (2023)
von: Kögel, Fabian, et al.
Veröffentlicht: (2023)
REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
von: Ouahidi, Yassine El, et al.
Veröffentlicht: (2025)
von: Ouahidi, Yassine El, et al.
Veröffentlicht: (2025)
Few and Fewer: Learning Better from Few Examples Using Fewer Base Classes
von: Lafargue, Raphael, et al.
Veröffentlicht: (2024)
von: Lafargue, Raphael, et al.
Veröffentlicht: (2024)
Unsupervised Adaptive Deep Learning Method For BCI Motor Imagery Decoding
von: Ouahidi, Yassine El, et al.
Veröffentlicht: (2024)
von: Ouahidi, Yassine El, et al.
Veröffentlicht: (2024)
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
von: Farid, Sualeha, et al.
Veröffentlicht: (2025)
von: Farid, Sualeha, et al.
Veröffentlicht: (2025)
From Sequence to Structure: Uncovering Substructure Reasoning in Transformers
von: Dai, Xinnan, et al.
Veröffentlicht: (2025)
von: Dai, Xinnan, et al.
Veröffentlicht: (2025)
TPTT: Transforming Pretrained Transformers into Titans
von: Furfaro, Fabien
Veröffentlicht: (2025)
von: Furfaro, Fabien
Veröffentlicht: (2025)
A Strong and Simple Deep Learning Baseline for BCI MI Decoding
von: Ouahidi, Yassine El, et al.
Veröffentlicht: (2023)
von: Ouahidi, Yassine El, et al.
Veröffentlicht: (2023)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
von: Xiao, Da, et al.
Veröffentlicht: (2025)
von: Xiao, Da, et al.
Veröffentlicht: (2025)
SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning
von: Nguyen, Bac, et al.
Veröffentlicht: (2024)
von: Nguyen, Bac, et al.
Veröffentlicht: (2024)
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
von: Ren, Qibing, et al.
Veröffentlicht: (2024)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2025)
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2025)
Can pre-trained Deep Learning models predict groove ratings?
von: Marmoret, Axel, et al.
Veröffentlicht: (2026)
von: Marmoret, Axel, et al.
Veröffentlicht: (2026)
Mitigating Misalignment Contagion by Steering with Implicit Traits
von: Chang, Maria, et al.
Veröffentlicht: (2026)
von: Chang, Maria, et al.
Veröffentlicht: (2026)
Order-Preserving GFlowNets
von: Chen, Yihang, et al.
Veröffentlicht: (2023)
von: Chen, Yihang, et al.
Veröffentlicht: (2023)
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
von: Joshi, Maithili, et al.
Veröffentlicht: (2025)
von: Joshi, Maithili, et al.
Veröffentlicht: (2025)
Semantic Containment as a Fundamental Property of Emergent Misalignment
von: Saxena, Rohan
Veröffentlicht: (2026)
von: Saxena, Rohan
Veröffentlicht: (2026)
Preemptive Detection and Correction of Misaligned Actions in LLM Agents
von: Fang, Haishuo, et al.
Veröffentlicht: (2024)
von: Fang, Haishuo, et al.
Veröffentlicht: (2024)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
LoR2C : Low-Rank Residual Connection Adaptation for Parameter-Efficient Fine-Tuning
von: Zhao, Jiancheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jiancheng, et al.
Veröffentlicht: (2025)
Residual Stream Duality in Modern Transformer Architectures
von: Zhang, Yifan
Veröffentlicht: (2026)
von: Zhang, Yifan
Veröffentlicht: (2026)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
von: Hu, Xuhao, et al.
Veröffentlicht: (2025)
An Investigation into Value Misalignment in LLM-Generated Texts for Cultural Heritage
von: Bu, Fan, et al.
Veröffentlicht: (2025)
von: Bu, Fan, et al.
Veröffentlicht: (2025)
NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured Noise
von: Xu, Zhi, et al.
Veröffentlicht: (2026)
von: Xu, Zhi, et al.
Veröffentlicht: (2026)
Transformer-based Causal Language Models Perform Clustering
von: Wu, Xinbo, et al.
Veröffentlicht: (2024)
von: Wu, Xinbo, et al.
Veröffentlicht: (2024)
Misaligned by Reward: Socially Undesirable Preferences in LLMs
von: Ghazaryan, Gayane, et al.
Veröffentlicht: (2026)
von: Ghazaryan, Gayane, et al.
Veröffentlicht: (2026)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models
von: Zhao, Juntu, et al.
Veröffentlicht: (2024)
von: Zhao, Juntu, et al.
Veröffentlicht: (2024)
StableMask: Refining Causal Masking in Decoder-only Transformer
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
ReCo: Reliable Causal Chain Reasoning via Structural Causal Recurrent Neural Networks
von: Xiong, Kai, et al.
Veröffentlicht: (2022)
von: Xiong, Kai, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training
von: Lys, Jonathan, et al.
Veröffentlicht: (2026) -
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding
von: Lys, Jonathan, et al.
Veröffentlicht: (2026) -
LLM meets Vision-Language Models for Zero-Shot One-Class Classification
von: Bendou, Yassir, et al.
Veröffentlicht: (2024) -
A Novel Benchmark for Few-Shot Semantic Segmentation in the Era of Foundation Models
von: Bensaid, Reda, et al.
Veröffentlicht: (2024) -
GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
von: Choudhary, Anand, et al.
Veröffentlicht: (2025)