Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
Fuente:
arXiv
Saved in:
| Main Authors: | Rai, Daking, Miller, Samuel, Moran, Kevin, Yao, Ziyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mechanistic Understanding of Language Models in Syntactic Code Completion
by: Miller, Samuel, et al.
Published: (2025)
by: Miller, Samuel, et al.
Published: (2025)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
Data-driven Circuit Discovery for Interpretability of Language Models
by: Rai, Daking, et al.
Published: (2026)
by: Rai, Daking, et al.
Published: (2026)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
by: Mamidanna, Siddarth, et al.
Published: (2025)
by: Mamidanna, Siddarth, et al.
Published: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
Can AI Assist in Olympiad Coding
by: Ren, Samuel
Published: (2025)
by: Ren, Samuel
Published: (2025)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025)
by: He, Kaifeng, et al.
Published: (2025)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
by: Abualazm, Raafat, et al.
Published: (2026)
by: Abualazm, Raafat, et al.
Published: (2026)
A Framework for Testing and Adapting REST APIs as LLM Tools
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
by: Bandlamudi, Jayachandu, et al.
Published: (2025)
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
by: Le, Nguyen-Khang, et al.
Published: (2025)
by: Le, Nguyen-Khang, et al.
Published: (2025)
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
by: Cartagena, Arnold, et al.
Published: (2026)
by: Cartagena, Arnold, et al.
Published: (2026)
The Appeal and Reality of Recycling LoRAs with Adaptive Merging
by: Liu, Haokun, et al.
Published: (2026)
by: Liu, Haokun, et al.
Published: (2026)
Engineering A Large Language Model From Scratch
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
by: Lawand, Daniel Angelo Esteves, et al.
Published: (2025)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026)
by: Weng, Haojun, et al.
Published: (2026)
Predicting 3D Rigid Body Dynamics with Deep Residual Network
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
Developer Challenges on Large Language Models: A Study of Stack Overflow and OpenAI Developer Forum Posts
by: Alam, Khairul, et al.
Published: (2024)
by: Alam, Khairul, et al.
Published: (2024)
Model-Driven Legacy System Modernization at Scale
by: Böhm, Tobias, et al.
Published: (2026)
by: Böhm, Tobias, et al.
Published: (2026)
CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research
by: Savenkov, Vladislav
Published: (2026)
by: Savenkov, Vladislav
Published: (2026)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
by: Kohl, Jens, et al.
Published: (2024)
by: Kohl, Jens, et al.
Published: (2024)
ContractBench: Can LLM Agents Preserve Observation Contracts?
by: Wang, Jicheng, et al.
Published: (2026)
by: Wang, Jicheng, et al.
Published: (2026)
FREYR: A Framework for Recognizing and Executing Your Requests
by: Gallotta, Roberto, et al.
Published: (2025)
by: Gallotta, Roberto, et al.
Published: (2025)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
by: Karpurapu, Shanthi, et al.
Published: (2024)
by: Karpurapu, Shanthi, et al.
Published: (2024)
Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering
by: Sharma, Krishna Kumaar
Published: (2025)
by: Sharma, Krishna Kumaar
Published: (2025)
The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot
by: Yeverechyahu, Doron, et al.
Published: (2024)
by: Yeverechyahu, Doron, et al.
Published: (2024)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition
by: Sellami, Khaled, et al.
Published: (2025)
by: Sellami, Khaled, et al.
Published: (2025)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
by: Liu, Zhuoyao, et al.
Published: (2026)
by: Liu, Zhuoyao, et al.
Published: (2026)
TCProF: Time-Complexity Prediction SSL Framework
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025)
by: Lim, Soohan, et al.
Published: (2025)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
by: Tran, Hung, et al.
Published: (2026)
by: Tran, Hung, et al.
Published: (2026)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
by: Zhu, Andrew, et al.
Published: (2024)
by: Zhu, Andrew, et al.
Published: (2024)
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
by: Sakizli, Furkan
Published: (2026)
by: Sakizli, Furkan
Published: (2026)
MicroRemed: Benchmarking LLMs in Microservices Remediation
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
by: Le, Nguyen-Khang, et al.
Published: (2025)
by: Le, Nguyen-Khang, et al.
Published: (2025)
Similar Items
-
Mechanistic Understanding of Language Models in Syntactic Code Completion
by: Miller, Samuel, et al.
Published: (2025) -
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs
by: Rai, Daking, et al.
Published: (2024) -
Data-driven Circuit Discovery for Interpretability of Language Models
by: Rai, Daking, et al.
Published: (2026) -
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
by: Mamidanna, Siddarth, et al.
Published: (2025) -
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)