Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
Fuente:
arXiv
Guardado en:
| Autores principales: | Shao, Shun, Wang, Binxu, Cohen, Shay B., Korhonen, Anna, Belinkov, Yonatan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
por: Hanna, Michael, et al.
Publicado: (2024)
por: Hanna, Michael, et al.
Publicado: (2024)
Iterative Multilingual Spectral Attribute Erasure
por: Shao, Shun, et al.
Publicado: (2025)
por: Shao, Shun, et al.
Publicado: (2025)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
por: Tutek, Martin, et al.
Publicado: (2025)
por: Tutek, Martin, et al.
Publicado: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
por: Simhi, Adi, et al.
Publicado: (2026)
por: Simhi, Adi, et al.
Publicado: (2026)
Spectral Editing of Activations for Large Language Model Alignment
por: Qiu, Yifu, et al.
Publicado: (2024)
por: Qiu, Yifu, et al.
Publicado: (2024)
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
por: Nikankin, Yaniv, et al.
Publicado: (2025)
por: Nikankin, Yaniv, et al.
Publicado: (2025)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
por: Rahamim, Adir, et al.
Publicado: (2023)
por: Rahamim, Adir, et al.
Publicado: (2023)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
por: Yu, Zeping, et al.
Publicado: (2025)
por: Yu, Zeping, et al.
Publicado: (2025)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
por: Qiu, Yifu, et al.
Publicado: (2025)
por: Qiu, Yifu, et al.
Publicado: (2025)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
por: Nikankin, Yaniv, et al.
Publicado: (2025)
por: Nikankin, Yaniv, et al.
Publicado: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
por: Ashuach, Tomer, et al.
Publicado: (2024)
por: Ashuach, Tomer, et al.
Publicado: (2024)
Think While You Write: Hypothesis Verification Promotes Faithful Knowledge-to-Text Generation
por: Qiu, Yifu, et al.
Publicado: (2023)
por: Qiu, Yifu, et al.
Publicado: (2023)
Position-aware Automatic Circuit Discovery
por: Haklay, Tal, et al.
Publicado: (2025)
por: Haklay, Tal, et al.
Publicado: (2025)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
por: Iskander, Shadi, et al.
Publicado: (2024)
por: Iskander, Shadi, et al.
Publicado: (2024)
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
por: Carmeli, Boaz, et al.
Publicado: (2024)
por: Carmeli, Boaz, et al.
Publicado: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
por: Arad, Dana, et al.
Publicado: (2023)
por: Arad, Dana, et al.
Publicado: (2023)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
por: Marks, Samuel, et al.
Publicado: (2024)
por: Marks, Samuel, et al.
Publicado: (2024)
Are formal and functional linguistic mechanisms dissociated in language models?
por: Hanna, Michael, et al.
Publicado: (2025)
por: Hanna, Michael, et al.
Publicado: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
por: Arad, Dana, et al.
Publicado: (2025)
por: Arad, Dana, et al.
Publicado: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
por: Itzhak, Itay, et al.
Publicado: (2025)
por: Itzhak, Itay, et al.
Publicado: (2025)
Structured RAG for Answering Aggregative Questions
por: Koshorek, Omri, et al.
Publicado: (2025)
por: Koshorek, Omri, et al.
Publicado: (2025)
DEPTH: Discourse Education through Pre-Training Hierarchically
por: Bamberger, Zachary, et al.
Publicado: (2024)
por: Bamberger, Zachary, et al.
Publicado: (2024)
Self-Improving World Modelling with Latent Actions
por: Qiu, Yifu, et al.
Publicado: (2026)
por: Qiu, Yifu, et al.
Publicado: (2026)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
por: Kaplan, Guy, et al.
Publicado: (2025)
por: Kaplan, Guy, et al.
Publicado: (2025)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
por: Arad, Dana, et al.
Publicado: (2025)
por: Arad, Dana, et al.
Publicado: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
por: Katz, Shahar, et al.
Publicado: (2024)
por: Katz, Shahar, et al.
Publicado: (2024)
Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
por: Fonseca, Marcio, et al.
Publicado: (2024)
por: Fonseca, Marcio, et al.
Publicado: (2024)
Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains
por: Fonseca, Marcio, et al.
Publicado: (2023)
por: Fonseca, Marcio, et al.
Publicado: (2023)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
por: Nikankin, Yaniv, et al.
Publicado: (2024)
por: Nikankin, Yaniv, et al.
Publicado: (2024)
Fast Forwarding Low-Rank Training
por: Rahamim, Adir, et al.
Publicado: (2024)
por: Rahamim, Adir, et al.
Publicado: (2024)
Will it Merge? On The Causes of Model Mergeability
por: Rahamim, Adir, et al.
Publicado: (2026)
por: Rahamim, Adir, et al.
Publicado: (2026)
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
por: Ventura, Mor, et al.
Publicado: (2025)
por: Ventura, Mor, et al.
Publicado: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
por: Simhi, Adi, et al.
Publicado: (2024)
por: Simhi, Adi, et al.
Publicado: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
por: Simhi, Adi, et al.
Publicado: (2024)
por: Simhi, Adi, et al.
Publicado: (2024)
Context-aware Prompt Tuning: Advancing In-Context Learning with Adversarial Methods
por: Blau, Tsachi, et al.
Publicado: (2024)
por: Blau, Tsachi, et al.
Publicado: (2024)
Modeling News Interactions and Influence for Financial Market Prediction
por: Wang, Mengyu, et al.
Publicado: (2024)
por: Wang, Mengyu, et al.
Publicado: (2024)
Silent Tokens, Loud Effects: Padding in LLMs
por: Himelstein, Rom, et al.
Publicado: (2025)
por: Himelstein, Rom, et al.
Publicado: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
por: Itzhak, Itay, et al.
Publicado: (2026)
por: Itzhak, Itay, et al.
Publicado: (2026)
Reasoning Models Know What's Important, and Encode It in Their Activations
por: Nikankin, Yaniv, et al.
Publicado: (2026)
por: Nikankin, Yaniv, et al.
Publicado: (2026)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
por: Nur'aini, Khumaisa, et al.
Publicado: (2026)
por: Nur'aini, Khumaisa, et al.
Publicado: (2026)
Ejemplares similares
-
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
por: Hanna, Michael, et al.
Publicado: (2024) -
Iterative Multilingual Spectral Attribute Erasure
por: Shao, Shun, et al.
Publicado: (2025) -
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
por: Tutek, Martin, et al.
Publicado: (2025) -
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
por: Simhi, Adi, et al.
Publicado: (2026) -
Spectral Editing of Activations for Large Language Model Alignment
por: Qiu, Yifu, et al.
Publicado: (2024)