Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Quirke, Philip, Oozeer, Narmeen, Bandi, Chaithanya, Abdullah, Amir, Hoelscher-Obermaier, Jason, Phillips, Jeff M., Greaves, Joshua, Neo, Clement, Lan, Michael, Barez, Fazl, Upadhyay, Shriyash |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Position: Require Frontier AI Labs To Release Small "Analog" Models
by: Upadhyay, Shriyash, et al.
Published: (2025)
by: Upadhyay, Shriyash, et al.
Published: (2025)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)
by: Ivanov, Georgi, et al.
Published: (2026)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Interpreting Learned Feedback Patterns in Large Language Models
by: Marks, Luke, et al.
Published: (2023)
by: Marks, Luke, et al.
Published: (2023)
Das lyrische Werk Antoni Langes
by: Hoelscher-Obermaier, Hans-Peter
Published: (2019)
by: Hoelscher-Obermaier, Hans-Peter
Published: (2019)
Andrzej Kuśniewicz' synkretistische Romanpoetik
by: Hoelscher-Obermaier, Hans-Peter
Published: (2019)
by: Hoelscher-Obermaier, Hans-Peter
Published: (2019)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
by: Harrasse, Abir, et al.
Published: (2025)
by: Harrasse, Abir, et al.
Published: (2025)
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
by: Harrasse, Abir, et al.
Published: (2024)
by: Harrasse, Abir, et al.
Published: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
Query Circuits: Explaining How Language Models Answer User Prompts
by: Wu, Tung-Yu, et al.
Published: (2025)
by: Wu, Tung-Yu, et al.
Published: (2025)
Activation Space Interventions Can Be Transferred Between Large Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Rethinking AI Cultural Alignment
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Token Taxes: mitigating AGI's economic risks
by: Irwin, Lucas, et al.
Published: (2026)
by: Irwin, Lucas, et al.
Published: (2026)
Large Language Models Relearn Removed Concepts
by: Lo, Michelle, et al.
Published: (2024)
by: Lo, Michelle, et al.
Published: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
REL: Working out is all you need
by: Simonds, Toby, et al.
Published: (2024)
by: Simonds, Toby, et al.
Published: (2024)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
by: Heindrich, Lovis, et al.
Published: (2025)
by: Heindrich, Lovis, et al.
Published: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Visualizing Neural Network Imagination
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent Systems
by: Peigne-Lefebvre, Pierre, et al.
Published: (2025)
by: Peigne-Lefebvre, Pierre, et al.
Published: (2025)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
by: Roytburg, Dani, et al.
Published: (2025)
by: Roytburg, Dani, et al.
Published: (2025)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Embodied AI: Emerging Risks and Opportunities for Policy Action
by: Perlo, Jared, et al.
Published: (2025)
by: Perlo, Jared, et al.
Published: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
Chain-of-Thought Hijacking
by: Zhao, Jianli, et al.
Published: (2025)
by: Zhao, Jianli, et al.
Published: (2025)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
by: Li, Changyi, et al.
Published: (2026)
by: Li, Changyi, et al.
Published: (2026)
Scaling sparse feature circuit finding for in-context learning
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
by: Kharlapenko, Dmitrii, et al.
Published: (2025)
Familienunternehmer als externe Beiräte
by: Obermaier, Otto W.
Published: (2020)
by: Obermaier, Otto W.
Published: (2020)
Similar Items
-
Position: Require Frontier AI Labs To Release Small "Analog" Models
by: Upadhyay, Shriyash, et al.
Published: (2025) -
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024) -
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025) -
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026) -
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)