Mechanistic Permutability: Match Features Across Layers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Balagansky, Nikita, Maksimov, Ian, Gavrilov, Daniil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
von: Balagansky, Nikita, et al.
Veröffentlicht: (2025)
von: Balagansky, Nikita, et al.
Veröffentlicht: (2025)
Learn Your Reference Model for Real Good Alignment
von: Gorbatovski, Alexey, et al.
Veröffentlicht: (2024)
von: Gorbatovski, Alexey, et al.
Veröffentlicht: (2024)
Next Embedding Prediction Makes World Models Stronger
von: Bredis, George, et al.
Veröffentlicht: (2026)
von: Bredis, George, et al.
Veröffentlicht: (2026)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
von: Kurochkin, Vadim, et al.
Veröffentlicht: (2025)
von: Kurochkin, Vadim, et al.
Veröffentlicht: (2025)
Diffusion Language Models Generation Can Be Halted Early
von: Vaina, Sofia Maria Lo Cicero, et al.
Veröffentlicht: (2023)
von: Vaina, Sofia Maria Lo Cicero, et al.
Veröffentlicht: (2023)
Teach Old SAEs New Domain Tricks with Boosting
von: Koriagin, Nikita, et al.
Veröffentlicht: (2025)
von: Koriagin, Nikita, et al.
Veröffentlicht: (2025)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
You Do Not Fully Utilize Transformer's Representation Capacity
von: Gerasimov, Gleb, et al.
Veröffentlicht: (2025)
von: Gerasimov, Gleb, et al.
Veröffentlicht: (2025)
Trust-Region Behavior Blending for On-Policy Distillation
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
Steering LLM Reasoning Through Bias-Only Adaptation
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
von: Sinii, Viacheslav, et al.
Veröffentlicht: (2025)
Revisiting Non-Acyclic GFlowNets in Discrete Environments
von: Morozov, Nikita, et al.
Veröffentlicht: (2025)
von: Morozov, Nikita, et al.
Veröffentlicht: (2025)
Guided Star-Shaped Masked Diffusion
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
Linear Transformers with Learnable Kernel Functions are Better In-Context Models
von: Aksenov, Yaroslav, et al.
Veröffentlicht: (2024)
von: Aksenov, Yaroslav, et al.
Veröffentlicht: (2024)
Learning Shortest Paths with Generative Flow Networks
von: Morozov, Nikita, et al.
Veröffentlicht: (2026)
von: Morozov, Nikita, et al.
Veröffentlicht: (2026)
gfnx: Fast and Scalable Library for Generative Flow Networks in JAX
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
von: Diatlova, Daria, et al.
Veröffentlicht: (2025)
von: Diatlova, Daria, et al.
Veröffentlicht: (2025)
Adversarial Schrödinger Bridge Matching
von: Gushchin, Nikita, et al.
Veröffentlicht: (2024)
von: Gushchin, Nikita, et al.
Veröffentlicht: (2024)
The Differences Between Direct Alignment Algorithms are a Blur
von: Gorbatovski, Alexey, et al.
Veröffentlicht: (2025)
von: Gorbatovski, Alexey, et al.
Veröffentlicht: (2025)
Evolution of SAE Features Across Layers in LLMs
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
von: Balcells, Daniel, et al.
Veröffentlicht: (2024)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
von: Bredis, George, et al.
Veröffentlicht: (2025)
von: Bredis, George, et al.
Veröffentlicht: (2025)
Inverse Bridge Matching Distillation
von: Gushchin, Nikita, et al.
Veröffentlicht: (2025)
von: Gushchin, Nikita, et al.
Veröffentlicht: (2025)
PPFS: Predictive Permutation Feature Selection
von: Hassan, Atif, et al.
Veröffentlicht: (2021)
von: Hassan, Atif, et al.
Veröffentlicht: (2021)
ESSA: Evolutionary Strategies for Scalable Alignment
von: Korotyshova, Daria, et al.
Veröffentlicht: (2025)
von: Korotyshova, Daria, et al.
Veröffentlicht: (2025)
Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates
von: Maksimov, Roman, et al.
Veröffentlicht: (2026)
von: Maksimov, Roman, et al.
Veröffentlicht: (2026)
Fair Feature Importance Scores via Feature Occlusion and Permutation
von: Little, Camille, et al.
Veröffentlicht: (2026)
von: Little, Camille, et al.
Veröffentlicht: (2026)
Trustworthy Feature Importance Avoids Unrestricted Permutations
von: Borgonovo, Emanuele, et al.
Veröffentlicht: (2026)
von: Borgonovo, Emanuele, et al.
Veröffentlicht: (2026)
Analysis of Linear Mode Connectivity via Permutation-Based Weight Matching: With Insights into Other Permutation Search Methods
von: Ito, Akira, et al.
Veröffentlicht: (2024)
von: Ito, Akira, et al.
Veröffentlicht: (2024)
Learning Unbiased Permutations via Flow Matching
von: Min, Yimeng, et al.
Veröffentlicht: (2026)
von: Min, Yimeng, et al.
Veröffentlicht: (2026)
Improving Generalization by Permutation Routing Across Model Copies
von: Kashiwamura, Shuhei, et al.
Veröffentlicht: (2026)
von: Kashiwamura, Shuhei, et al.
Veröffentlicht: (2026)
Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization
von: Gritsaev, Timofei, et al.
Veröffentlicht: (2024)
von: Gritsaev, Timofei, et al.
Veröffentlicht: (2024)
Generative Flow Networks as Entropy-Regularized RL
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
Unlocking the Duality between Flow and Field Matching
von: Shlenskii, Daniil, et al.
Veröffentlicht: (2026)
von: Shlenskii, Daniil, et al.
Veröffentlicht: (2026)
On the Equivalence of Optimal Transport Problem and Action Matching with Optimal Vector Fields
von: Kornilov, Nikita, et al.
Veröffentlicht: (2025)
von: Kornilov, Nikita, et al.
Veröffentlicht: (2025)
UPath: Universal Planner Across Topological Heterogeneity For Grid-Based Pathfinding
von: Ananikian, Aleksandr, et al.
Veröffentlicht: (2026)
von: Ananikian, Aleksandr, et al.
Veröffentlicht: (2026)
Tracking Equivalent Mechanistic Interpretations Across Neural Networks
von: Sun, Alan, et al.
Veröffentlicht: (2026)
von: Sun, Alan, et al.
Veröffentlicht: (2026)
AI Methods for Permutation Circuit Synthesis Across Generic Topologies
von: Villar, Victor, et al.
Veröffentlicht: (2025)
von: Villar, Victor, et al.
Veröffentlicht: (2025)
Adaptive Set-Mass Calibration with Conformal Prediction
von: Kazantsev, Daniil, et al.
Veröffentlicht: (2025)
von: Kazantsev, Daniil, et al.
Veröffentlicht: (2025)
Light and Optimal Schrödinger Bridge Matching
von: Gushchin, Nikita, et al.
Veröffentlicht: (2024)
von: Gushchin, Nikita, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025) -
Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy
von: Balagansky, Nikita, et al.
Veröffentlicht: (2025) -
Learn Your Reference Model for Real Good Alignment
von: Gorbatovski, Alexey, et al.
Veröffentlicht: (2024) -
Next Embedding Prediction Makes World Models Stronger
von: Bredis, George, et al.
Veröffentlicht: (2026) -
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
von: Kurochkin, Vadim, et al.
Veröffentlicht: (2025)