Studying Cross-cluster Modularity in Neural Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Golechha, Satvik, Chaudhary, Maheep, Velja, Joan, Abate, Alessandro, Schoots, Nandi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training Neural Networks for Modularity aids Interpretability
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Progress Measures for Grokking on Real-world Tasks
von: Golechha, Satvik
Veröffentlicht: (2024)
von: Golechha, Satvik
Veröffentlicht: (2024)
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
von: Golechha, Satvik, et al.
Veröffentlicht: (2024)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
von: Golechha, Satvik, et al.
Veröffentlicht: (2025)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
von: Chaudhary, Siddharth, et al.
Veröffentlicht: (2025)
von: Chaudhary, Siddharth, et al.
Veröffentlicht: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
von: Chaudhary, Maheep
Veröffentlicht: (2026)
von: Chaudhary, Maheep
Veröffentlicht: (2026)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Relating Piecewise Linear Kolmogorov Arnold Networks to ReLU Networks
von: Schoots, Nandi, et al.
Veröffentlicht: (2025)
von: Schoots, Nandi, et al.
Veröffentlicht: (2025)
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
Building Better Deception Probes Using Targeted Instruction Pairs
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
von: Natarajan, Vikram, et al.
Veröffentlicht: (2026)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2024)
Punctuation and Predicates in Language Models
von: Chauhan, Sonakshi, et al.
Veröffentlicht: (2025)
von: Chauhan, Sonakshi, et al.
Veröffentlicht: (2025)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
von: Lidayan, Aly, et al.
Veröffentlicht: (2025)
Who's the Evil Twin? Differential Auditing for Undesired Behavior
von: Balappanawar, Ishwar, et al.
Veröffentlicht: (2025)
von: Balappanawar, Ishwar, et al.
Veröffentlicht: (2025)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
von: Egbuna, Nathan, et al.
Veröffentlicht: (2025)
von: Egbuna, Nathan, et al.
Veröffentlicht: (2025)
The Propensity for Density in Feed-forward Models
von: Schoots, Nandi, et al.
Veröffentlicht: (2024)
von: Schoots, Nandi, et al.
Veröffentlicht: (2024)
Neural Proofs for Sound Verification and Control of Complex Systems
von: Abate, Alessandro
Veröffentlicht: (2025)
von: Abate, Alessandro
Veröffentlicht: (2025)
Extending Activation Steering to Broad Skills and Multiple Behaviours
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
von: van der Weij, Teun, et al.
Veröffentlicht: (2024)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2024)
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2024)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
von: Skalse, Joar, et al.
Veröffentlicht: (2024)
Soft Contamination Means Benchmarks Test Shallow Generalization
von: Spiesberger, Ari, et al.
Veröffentlicht: (2026)
von: Spiesberger, Ari, et al.
Veröffentlicht: (2026)
Modular Boundaries in Recurrent Neural Networks
von: Tanner, Jacob, et al.
Veröffentlicht: (2023)
von: Tanner, Jacob, et al.
Veröffentlicht: (2023)
Dynamic Vocabulary Pruning in Early-Exit LLMs
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
von: Vincenti, Jort, et al.
Veröffentlicht: (2024)
PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
von: Vuddanti, Sri Vatsa, et al.
Veröffentlicht: (2025)
von: Vuddanti, Sri Vatsa, et al.
Veröffentlicht: (2025)
Broken Chains: The Cost of Incomplete Reasoning in LLMs
von: Su, Ian, et al.
Veröffentlicht: (2026)
von: Su, Ian, et al.
Veröffentlicht: (2026)
Zero-Shot Instruction Following in RL via Structured LTL Representations
von: Giuri, Mattia, et al.
Veröffentlicht: (2025)
von: Giuri, Mattia, et al.
Veröffentlicht: (2025)
Efficient Solution and Learning of Robust Factored MDPs
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
Weight space Detection of Backdoors in LoRA Adapters
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
von: Merenciano, David Puertolas, et al.
Veröffentlicht: (2026)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
Temporal-Difference Variational Continual Learning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
von: More, Abhishek, et al.
Veröffentlicht: (2025)
von: More, Abhishek, et al.
Veröffentlicht: (2025)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
von: Nunez, Jeanmely Rojas, et al.
Veröffentlicht: (2026)
Networked Communication for Decentralised Agents in Mean-Field Games
von: Benjamin, Patrick, et al.
Veröffentlicht: (2023)
von: Benjamin, Patrick, et al.
Veröffentlicht: (2023)
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
von: Griffin, Charlie, et al.
Veröffentlicht: (2024)
Zero-Shot Instruction Following in RL via Structured LTL Representations
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2026)
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2026)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
von: Swaroop, Anand, et al.
Veröffentlicht: (2025)
von: Swaroop, Anand, et al.
Veröffentlicht: (2025)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
von: Patel, Dev, et al.
Veröffentlicht: (2025)
von: Patel, Dev, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Training Neural Networks for Modularity aids Interpretability
von: Golechha, Satvik, et al.
Veröffentlicht: (2024) -
Progress Measures for Grokking on Real-world Tasks
von: Golechha, Satvik
Veröffentlicht: (2024) -
Challenges in Mechanistically Interpreting Model Representations
von: Golechha, Satvik, et al.
Veröffentlicht: (2024) -
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
von: Golechha, Satvik, et al.
Veröffentlicht: (2025) -
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
von: Chaudhary, Siddharth, et al.
Veröffentlicht: (2025)