BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Mondorf, Philipp, Wang, Mingyang, Gerstner, Sebastian, Hakimi, Ahmad Dawar, Liu, Yihong, Veloso, Leonor, Zhou, Shijia, Schütze, Hinrich, Plank, Barbara |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
by: Nikankin, Yaniv, et al.
Published: (2025)
by: Nikankin, Yaniv, et al.
Published: (2025)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
by: Hakimi, Ahmad Dawar, et al.
Published: (2025)
by: Hakimi, Ahmad Dawar, et al.
Published: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
by: Hakimi, Ahmad Dawar, et al.
Published: (2026)
by: Hakimi, Ahmad Dawar, et al.
Published: (2026)
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
by: Veloso, Leonor, et al.
Published: (2026)
by: Veloso, Leonor, et al.
Published: (2026)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
by: Mondorf, Philipp, et al.
Published: (2025)
by: Mondorf, Philipp, et al.
Published: (2025)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
by: Gerstner, Sebastian, et al.
Published: (2026)
by: Gerstner, Sebastian, et al.
Published: (2026)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
by: Gerstner, Sebastian, et al.
Published: (2025)
by: Gerstner, Sebastian, et al.
Published: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
On Relation-Specific Neurons in Large Language Models
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
SLAyiNG: Towards Queer Language Processing
by: Veloso, Leonor, et al.
Published: (2025)
by: Veloso, Leonor, et al.
Published: (2025)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
by: Orth, Jasmin, et al.
Published: (2025)
by: Orth, Jasmin, et al.
Published: (2025)
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
by: Rabern, Brian, et al.
Published: (2026)
by: Rabern, Brian, et al.
Published: (2026)
Add Noise, Tasks, or Layers? MaiNLP at the VarDial 2025 Shared Task on Norwegian Dialectal Slot and Intent Detection
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
by: Liu, Yihong, et al.
Published: (2023)
by: Liu, Yihong, et al.
Published: (2023)
On the Entity-Level Alignment in Crosslingual Consistency
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
MaiNLP at SemEval-2024 Task 1: Analyzing Source Language Selection in Cross-Lingual Textual Relatedness
by: Zhou, Shijia, et al.
Published: (2024)
by: Zhou, Shijia, et al.
Published: (2024)
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
by: Bertolazzi, Leonardo, et al.
Published: (2025)
by: Bertolazzi, Leonardo, et al.
Published: (2025)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
by: Cheng, Xinyuan, et al.
Published: (2026)
by: Cheng, Xinyuan, et al.
Published: (2026)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
How Programming Concepts and Neurons Are Shared in Code Language Models
by: Kargaran, Amir Hossein, et al.
Published: (2025)
by: Kargaran, Amir Hossein, et al.
Published: (2025)
Tracing Multilingual Factual Knowledge Acquisition in Pretraining
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
SYNTHEVAL: Hybrid Behavioral Testing of NLP Models with Synthetic CheckLists
by: Zhao, Raoyuan, et al.
Published: (2024)
by: Zhao, Raoyuan, et al.
Published: (2024)
LangSAMP: Language-Script Aware Multilingual Pretraining
by: Liu, Yihong, et al.
Published: (2024)
by: Liu, Yihong, et al.
Published: (2024)
Exploring the Role of Transliteration in In-Context Learning for Low-resource Languages Written in Non-Latin Scripts
by: Ma, Chunlan, et al.
Published: (2024)
by: Ma, Chunlan, et al.
Published: (2024)
MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
Relational Linearity is a Predictor of Hallucinations
by: Lu, Yuetian, et al.
Published: (2026)
by: Lu, Yuetian, et al.
Published: (2026)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
HYPEROFA: Expanding LLM Vocabulary to New Languages via Hypernetwork-Based Embedding Initialization
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
by: Feng, Qi, et al.
Published: (2025)
by: Feng, Qi, et al.
Published: (2025)
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
by: Veitsman, Yana, et al.
Published: (2026)
by: Veitsman, Yana, et al.
Published: (2026)
Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment
by: Xhelili, Orgest, et al.
Published: (2024)
by: Xhelili, Orgest, et al.
Published: (2024)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
by: Chen, Qiqi, et al.
Published: (2024)
by: Chen, Qiqi, et al.
Published: (2024)
NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis
by: Wang, Mingyang, et al.
Published: (2023)
by: Wang, Mingyang, et al.
Published: (2023)
MoSECroT: Model Stitching with Static Word Embeddings for Crosslingual Zero-shot Transfer
by: Ye, Haotian, et al.
Published: (2024)
by: Ye, Haotian, et al.
Published: (2024)
Similar Items
-
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
by: Nikankin, Yaniv, et al.
Published: (2025) -
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
by: Arad, Dana, et al.
Published: (2025) -
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models
by: Hakimi, Ahmad Dawar, et al.
Published: (2025) -
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
by: Mondorf, Philipp, et al.
Published: (2024) -
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
by: Hakimi, Ahmad Dawar, et al.
Published: (2026)