Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
Fuente:
arXiv
Saved in:
| Main Authors: | Harrasse, Abir, Draye, Florent, Pandey, Punya Syon, Jin, Zhijing, Schölkopf, Bernhard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
by: Draye, Florent, et al.
Published: (2026)
by: Draye, Florent, et al.
Published: (2026)
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026)
by: Pandey, Punya Syon, et al.
Published: (2026)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
by: Miller, Moritz, et al.
Published: (2026)
by: Miller, Moritz, et al.
Published: (2026)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024)
by: Ceraolo, Roberto, et al.
Published: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
by: Kassem, Aly M., et al.
Published: (2025)
by: Kassem, Aly M., et al.
Published: (2025)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
by: Ortu, Francesco, et al.
Published: (2026)
by: Ortu, Francesco, et al.
Published: (2026)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
by: Ortu, Francesco, et al.
Published: (2024)
by: Ortu, Francesco, et al.
Published: (2024)
Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
by: Harrasse, Abir, et al.
Published: (2024)
by: Harrasse, Abir, et al.
Published: (2024)
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning
by: Cui, Shaobo, et al.
Published: (2024)
by: Cui, Shaobo, et al.
Published: (2024)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
by: Lyu, Zhiheng, et al.
Published: (2024)
by: Lyu, Zhiheng, et al.
Published: (2024)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
by: Vennemeyer, Daniel, et al.
Published: (2026)
by: Vennemeyer, Daniel, et al.
Published: (2026)
A diverse Multilingual News Headlines Dataset from around the World
by: Leeb, Felix, et al.
Published: (2024)
by: Leeb, Felix, et al.
Published: (2024)
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal
by: Aggarwal, Vaibhav, et al.
Published: (2025)
by: Aggarwal, Vaibhav, et al.
Published: (2025)
Language Model Alignment in Multilingual Trolley Problems
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
by: Piatti, Giorgio, et al.
Published: (2024)
by: Piatti, Giorgio, et al.
Published: (2024)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
by: Jenny, David F., et al.
Published: (2023)
by: Jenny, David F., et al.
Published: (2023)
Are Language Models Consequentialist or Deontological Moral Reasoners?
by: Samway, Keenan, et al.
Published: (2025)
by: Samway, Keenan, et al.
Published: (2025)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
by: Backmann, Steffen, et al.
Published: (2025)
by: Backmann, Steffen, et al.
Published: (2025)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
When Do Language Models Endorse Limitations on Human Rights Principles?
by: Samway, Keenan, et al.
Published: (2026)
by: Samway, Keenan, et al.
Published: (2026)
Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models
by: Damianos, Dimitrios, et al.
Published: (2026)
by: Damianos, Dimitrios, et al.
Published: (2026)
Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution
by: Chen, Qinhao, et al.
Published: (2026)
by: Chen, Qinhao, et al.
Published: (2026)
Causal Responsibility Attribution for Human-AI Collaboration
by: Qi, Yahang, et al.
Published: (2024)
by: Qi, Yahang, et al.
Published: (2024)
The Curious Case of Curiosity across Human Cultures and LLMs
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
CausalCite: A Causal Formulation of Paper Citations
by: Kumar, Ishan, et al.
Published: (2023)
by: Kumar, Ishan, et al.
Published: (2023)
Can Theoretical Physics Research Benefit from Language Agents?
by: Lu, Sirui, et al.
Published: (2025)
by: Lu, Sirui, et al.
Published: (2025)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Implicit Personalization in Language Models: A Systematic Study
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Causality for Natural Language Processing
by: Jin, Zhijing
Published: (2025)
by: Jin, Zhijing
Published: (2025)
Nullpointer at CheckThat! 2024: Identifying Subjectivity from Multilingual Text Sequence
by: Biswas, Md. Rafiul, et al.
Published: (2024)
by: Biswas, Md. Rafiul, et al.
Published: (2024)
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Intrinsically Interpretable Attention via Sparse Post-Training
by: Draye, Florent, et al.
Published: (2025)
by: Draye, Florent, et al.
Published: (2025)
LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback
by: Lai, Wen, et al.
Published: (2024)
by: Lai, Wen, et al.
Published: (2024)
Similar Items
-
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
by: Draye, Florent, et al.
Published: (2026) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026) -
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025) -
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
by: Pandey, Punya Syon, et al.
Published: (2025) -
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
by: Miller, Moritz, et al.
Published: (2026)