Query Circuits: Explaining How Language Models Answer User Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Tung-Yu, Barez, Fazl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Large Language Models Relearn Removed Concepts
von: Lo, Michelle, et al.
Veröffentlicht: (2024)
von: Lo, Michelle, et al.
Veröffentlicht: (2024)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
Rethinking AI Cultural Alignment
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Visualizing Neural Network Imagination
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
Chain-of-Thought Hijacking
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
von: Li, Changyi, et al.
Veröffentlicht: (2026)
von: Li, Changyi, et al.
Veröffentlicht: (2026)
Same Answer, Different Representations: Hidden instability in VLMs
von: Wani, Farooq Ahmad, et al.
Veröffentlicht: (2026)
von: Wani, Farooq Ahmad, et al.
Veröffentlicht: (2026)
Embodied AI: Emerging Risks and Opportunities for Policy Action
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
von: Perlo, Jared, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Token Taxes: mitigating AGI's economic risks
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
Logic Query of Thoughts: Guiding Large Language Models to Answer Complex Logic Queries with Knowledge Graphs
von: Liu, Lihui, et al.
Veröffentlicht: (2024)
von: Liu, Lihui, et al.
Veröffentlicht: (2024)
Word Importance Explains How Prompts Affect Language Model Outputs
von: Hackmann, Stefan, et al.
Veröffentlicht: (2024)
von: Hackmann, Stefan, et al.
Veröffentlicht: (2024)
CircuitFormer: A Circuit Language Model for Analog Topology Design from Natural Language Prompt
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2026)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2026)
Formally Explaining Decision Tree Models with Answer Set Programming
von: Takemura, Akihiro, et al.
Veröffentlicht: (2026)
von: Takemura, Akihiro, et al.
Veröffentlicht: (2026)
Understanding LLM Scientific Reasoning through Promptings and Model's Explanation on the Answers
von: Rueda, Alice, et al.
Veröffentlicht: (2025)
von: Rueda, Alice, et al.
Veröffentlicht: (2025)
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
von: Mishra, Manit, et al.
Veröffentlicht: (2024)
Finding Interpretable Prompt-Specific Circuits in Language Models
von: Franco, Gabriel, et al.
Veröffentlicht: (2026)
von: Franco, Gabriel, et al.
Veröffentlicht: (2026)
Sandwich Reasoning: An Answer-Reasoning-Answer Approach for Low-Latency Query Correction
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
von: Raval, Shivam, et al.
Veröffentlicht: (2026)
von: Raval, Shivam, et al.
Veröffentlicht: (2026)
Shapley Revisited: Tractable Responsibility Measures for Query Answers
von: Bienvenu, Meghyn, et al.
Veröffentlicht: (2025)
von: Bienvenu, Meghyn, et al.
Veröffentlicht: (2025)
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
von: Liu, Tong, et al.
Veröffentlicht: (2024)
von: Liu, Tong, et al.
Veröffentlicht: (2024)
U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language Models
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
CLMASP: Coupling Large Language Models with Answer Set Programming for Robotic Task Planning
von: Lin, Xinrui, et al.
Veröffentlicht: (2024)
von: Lin, Xinrui, et al.
Veröffentlicht: (2024)
Explain Before You Answer: A Survey on Compositional Visual Reasoning
von: Ke, Fucai, et al.
Veröffentlicht: (2025)
von: Ke, Fucai, et al.
Veröffentlicht: (2025)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
von: Denison, Carson, et al.
Veröffentlicht: (2024)
von: Denison, Carson, et al.
Veröffentlicht: (2024)
Does This Summary Answer My Question? Modeling Query-Focused Summary Readers with Rational Speech Acts
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2024)
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2024)
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions
von: Bašaragin, Bojana, et al.
Veröffentlicht: (2024)
von: Bašaragin, Bojana, et al.
Veröffentlicht: (2024)
Susceptibility of Large Language Models to User-Driven Factors in Medical Queries
von: Lim, Kyung Ho, et al.
Veröffentlicht: (2025)
von: Lim, Kyung Ho, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023) -
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025) -
Large Language Models Relearn Removed Concepts
von: Lo, Michelle, et al.
Veröffentlicht: (2024) -
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025) -
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)