CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
Fuente:
arXiv
Salvato in:
| Autori principali: | Draye, Florent, Harrasse, Abir, Palit, Vedant, Wu, Tung-Yu, Liu, Jiarui, Pandey, Punya Syon, Wu, Roderick, Zhang, Terry Jingchen, Jin, Zhijing, Schölkopf, Bernhard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
di: Miller, Moritz, et al.
Pubblicazione: (2026)
di: Miller, Moritz, et al.
Pubblicazione: (2026)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
di: Liu, Xinge, et al.
Pubblicazione: (2026)
di: Liu, Xinge, et al.
Pubblicazione: (2026)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
di: Ortu, Francesco, et al.
Pubblicazione: (2026)
di: Ortu, Francesco, et al.
Pubblicazione: (2026)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Causal Responsibility Attribution for Human-AI Collaboration
di: Qi, Yahang, et al.
Pubblicazione: (2024)
di: Qi, Yahang, et al.
Pubblicazione: (2024)
Intrinsically Interpretable Attention via Sparse Post-Training
di: Draye, Florent, et al.
Pubblicazione: (2025)
di: Draye, Florent, et al.
Pubblicazione: (2025)
Can Theoretical Physics Research Benefit from Language Agents?
di: Lu, Sirui, et al.
Pubblicazione: (2025)
di: Lu, Sirui, et al.
Pubblicazione: (2025)
Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks
di: Palit, Vedant
Pubblicazione: (2025)
di: Palit, Vedant
Pubblicazione: (2025)
Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
di: Harrasse, Abir, et al.
Pubblicazione: (2024)
di: Harrasse, Abir, et al.
Pubblicazione: (2024)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2026)
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2026)
Causality can systematically address the monsters under the bench(marks)
di: Leeb, Felix, et al.
Pubblicazione: (2025)
di: Leeb, Felix, et al.
Pubblicazione: (2025)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
di: Jenny, David F., et al.
Pubblicazione: (2023)
di: Jenny, David F., et al.
Pubblicazione: (2023)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
di: Kassem, Aly M., et al.
Pubblicazione: (2025)
di: Kassem, Aly M., et al.
Pubblicazione: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
di: Simko, Samuel, et al.
Pubblicazione: (2025)
di: Simko, Samuel, et al.
Pubblicazione: (2025)
The Odyssey of Commonsense Causality: From Foundational Benchmarks to Cutting-Edge Reasoning
di: Cui, Shaobo, et al.
Pubblicazione: (2024)
di: Cui, Shaobo, et al.
Pubblicazione: (2024)
Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution
di: Chen, Qinhao, et al.
Pubblicazione: (2026)
di: Chen, Qinhao, et al.
Pubblicazione: (2026)
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
di: Hossain, Saad, et al.
Pubblicazione: (2026)
di: Hossain, Saad, et al.
Pubblicazione: (2026)
DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal
di: Aggarwal, Vaibhav, et al.
Pubblicazione: (2025)
di: Aggarwal, Vaibhav, et al.
Pubblicazione: (2025)
Can Large Language Models Infer Causation from Correlation?
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Are LLMs Good Safety Agents or a Propaganda Engine?
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
Implicit Personalization in Language Models: A Systematic Study
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
di: Chatzoudis, Gerasimos, et al.
Pubblicazione: (2026)
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
di: Kim, Su-Hyeon, et al.
Pubblicazione: (2026)
di: Kim, Su-Hyeon, et al.
Pubblicazione: (2026)
Activation Space Interventions Can Be Transferred Between Large Language Models
di: Oozeer, Narmeen, et al.
Pubblicazione: (2025)
di: Oozeer, Narmeen, et al.
Pubblicazione: (2025)
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
Orthogonal Finetuning Made Scalable
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
di: Raval, Shivam, et al.
Pubblicazione: (2026)
di: Raval, Shivam, et al.
Pubblicazione: (2026)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
Scalable On-the-fly Transcoding for Adaptive Streaming of Dynamic Point Clouds
di: Rudolph, Michael, et al.
Pubblicazione: (2026)
di: Rudolph, Michael, et al.
Pubblicazione: (2026)
On CLT and non-CLT groups
di: Tărnăuceanu, Marius
Pubblicazione: (2024)
di: Tărnăuceanu, Marius
Pubblicazione: (2024)
Decomposing and Measuring Evaluation Awareness
di: Li, Changling, et al.
Pubblicazione: (2026)
di: Li, Changling, et al.
Pubblicazione: (2026)
Forging a link: How the new Chinese Company Law shapes directors' liabilities and duties towards creditors
di: Jingchen Zhao, et al.
Pubblicazione: (2026)
di: Jingchen Zhao, et al.
Pubblicazione: (2026)
Protein Circuit Tracing via Cross-layer Transcoders
di: Tsui, Darin, et al.
Pubblicazione: (2026)
di: Tsui, Darin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
di: Harrasse, Abir, et al.
Pubblicazione: (2025) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026) -
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025) -
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
di: Miller, Moritz, et al.
Pubblicazione: (2026) -
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)