Does Editing Provide Evidence for Localization?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zihao, Veitch, Victor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
by: Li, Yin
Published: (2025)
by: Li, Yin
Published: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
by: Zhang, Xue
Published: (2025)
by: Zhang, Xue
Published: (2025)
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
by: Zhao, Zhilong, et al.
Published: (2025)
by: Zhao, Zhilong, et al.
Published: (2025)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
by: Wang, Zhixiang
Published: (2025)
by: Wang, Zhixiang
Published: (2025)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
by: Henry, James
Published: (2026)
by: Henry, James
Published: (2026)
Measuring Intent Comprehension in LLMs
by: Kunievsky, Nadav, et al.
Published: (2025)
by: Kunievsky, Nadav, et al.
Published: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
by: Jia, Xiao
Published: (2026)
by: Jia, Xiao
Published: (2026)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
by: Platt, Nolan, et al.
Published: (2025)
by: Platt, Nolan, et al.
Published: (2025)
Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning
by: Chanda, Prateek, et al.
Published: (2025)
by: Chanda, Prateek, et al.
Published: (2025)
Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
by: Chen, Jingwei
Published: (2025)
by: Chen, Jingwei
Published: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
by: Fan, Jingxing, et al.
Published: (2025)
by: Fan, Jingxing, et al.
Published: (2025)
DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
by: Li, Yangyang
Published: (2025)
by: Li, Yangyang
Published: (2025)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
by: Kaiser, Daniel, et al.
Published: (2026)
by: Kaiser, Daniel, et al.
Published: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
by: Alpay, Faruk, et al.
Published: (2025)
by: Alpay, Faruk, et al.
Published: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
by: Kim, Heejun, et al.
Published: (2026)
by: Kim, Heejun, et al.
Published: (2026)
Integrating External Tools with Large Language Models to Improve Accuracy
by: Niketan, Nripesh, et al.
Published: (2025)
by: Niketan, Nripesh, et al.
Published: (2025)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
by: Pizzo, David Alejandro Trejo
Published: (2026)
by: Pizzo, David Alejandro Trejo
Published: (2026)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
by: Das, Sourav
Published: (2026)
by: Das, Sourav
Published: (2026)
TSDS: Data Selection for Task-Specific Model Finetuning
by: Liu, Zifan, et al.
Published: (2024)
by: Liu, Zifan, et al.
Published: (2024)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
by: Garg, Saloni, et al.
Published: (2026)
by: Garg, Saloni, et al.
Published: (2026)
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
ProactBench: Beyond What The User Asked For
by: Harfi, Sepehr, et al.
Published: (2026)
by: Harfi, Sepehr, et al.
Published: (2026)
Hopscotch: Discovering and Skipping Redundancies in Language Models
by: Eyceoz, Mustafa, et al.
Published: (2025)
by: Eyceoz, Mustafa, et al.
Published: (2025)
Aletheia: Quantifying Cognitive Conviction in Reasoning Models via Regularized Inverse Confusion Matrix
by: Fu, Fanzhe
Published: (2026)
by: Fu, Fanzhe
Published: (2026)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
by: Bajpai, Ashutosh, et al.
Published: (2026)
by: Bajpai, Ashutosh, et al.
Published: (2026)
Beyond Subtokens: A Rich Character Embedding for Low-resource and Morphologically Complex Languages
by: Schneider, Felix, et al.
Published: (2026)
by: Schneider, Felix, et al.
Published: (2026)
On the Compatibility of Generative AI and Generative Linguistics
by: Portelance, Eva, et al.
Published: (2024)
by: Portelance, Eva, et al.
Published: (2024)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis
by: Zanbaghi, Shahin, et al.
Published: (2025)
by: Zanbaghi, Shahin, et al.
Published: (2025)
Improving Commonsense Bias Classification by Mitigating the Influence of Demographic Terms
by: Lee, JinKyu, et al.
Published: (2024)
by: Lee, JinKyu, et al.
Published: (2024)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026)
by: Mahale, Ajay Pravin
Published: (2026)
Towards Ontology-Enhanced Representation Learning for Large Language Models
by: Ronzano, Francesco, et al.
Published: (2024)
by: Ronzano, Francesco, et al.
Published: (2024)
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
J6: Jacobian-Driven Role Attribution for Multi-Objective Prompt Optimization in LLMs
by: Wu, Yao
Published: (2025)
by: Wu, Yao
Published: (2025)
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
by: Juzek, Tom S., et al.
Published: (2025)
by: Juzek, Tom S., et al.
Published: (2025)
Similar Items
-
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
by: Li, Yin
Published: (2025) -
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
by: Zhang, Xue
Published: (2025) -
A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
by: Zhao, Zhilong, et al.
Published: (2025) -
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
by: Wang, Zhixiang
Published: (2025) -
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
by: Borobia, Hector, et al.
Published: (2026)