AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deiseroth, Björn, Deb, Mayukh, Weinbach, Samuel, Brack, Manuel, Schramowski, Patrick, Kersting, Kristian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
von: Sztwiertnia, Sebastian, et al.
Veröffentlicht: (2025)
von: Sztwiertnia, Sebastian, et al.
Veröffentlicht: (2025)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
von: Höth, Max Henning, et al.
Veröffentlicht: (2026)
von: Höth, Max Henning, et al.
Veröffentlicht: (2026)
Core Tokensets for Data-efficient Sequential Training of Transformers
von: Paul, Subarnaduti, et al.
Veröffentlicht: (2024)
von: Paul, Subarnaduti, et al.
Veröffentlicht: (2024)
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
von: Helff, Lukas, et al.
Veröffentlicht: (2024)
von: Helff, Lukas, et al.
Veröffentlicht: (2024)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
von: Härle, Ruben, et al.
Veröffentlicht: (2024)
von: Härle, Ruben, et al.
Veröffentlicht: (2024)
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
von: Struppek, Lukas, et al.
Veröffentlicht: (2022)
von: Struppek, Lukas, et al.
Veröffentlicht: (2022)
DeiSAM: Segment Anything with Deictic Prompting
von: Shindo, Hikaru, et al.
Veröffentlicht: (2024)
von: Shindo, Hikaru, et al.
Veröffentlicht: (2024)
Measuring and Guiding Monosemanticity
von: Härle, Ruben, et al.
Veröffentlicht: (2025)
von: Härle, Ruben, et al.
Veröffentlicht: (2025)
CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
von: Hegde, Niharika, et al.
Veröffentlicht: (2025)
von: Hegde, Niharika, et al.
Veröffentlicht: (2025)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
von: Brack, Manuel, et al.
Veröffentlicht: (2025)
von: Brack, Manuel, et al.
Veröffentlicht: (2025)
LEDITS++: Limitless Image Editing using Text-to-Image Models
von: Brack, Manuel, et al.
Veröffentlicht: (2023)
von: Brack, Manuel, et al.
Veröffentlicht: (2023)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
von: Deiseroth, Björn, et al.
Veröffentlicht: (2025)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2025)
A Typology for Exploring the Mitigation of Shortcut Behavior
von: Friedrich, Felix, et al.
Veröffentlicht: (2022)
von: Friedrich, Felix, et al.
Veröffentlicht: (2022)
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
von: Helff, Lukas, et al.
Veröffentlicht: (2025)
von: Helff, Lukas, et al.
Veröffentlicht: (2025)
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
von: Shindo, Hikaru, et al.
Veröffentlicht: (2026)
von: Shindo, Hikaru, et al.
Veröffentlicht: (2026)
ART: Adaptive Relation Tuning for Generalized Relation Prediction
von: Sudhakaran, Gopika, et al.
Veröffentlicht: (2025)
von: Sudhakaran, Gopika, et al.
Veröffentlicht: (2025)
Learning by Self-Explaining
von: Stammer, Wolfgang, et al.
Veröffentlicht: (2023)
von: Stammer, Wolfgang, et al.
Veröffentlicht: (2023)
Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization
von: Deiseroth, Björn, et al.
Veröffentlicht: (2023)
von: Deiseroth, Björn, et al.
Veröffentlicht: (2023)
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
von: Friedrich, Felix, et al.
Veröffentlicht: (2025)
von: Friedrich, Felix, et al.
Veröffentlicht: (2025)
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
von: Neitemeier, Pit, et al.
Veröffentlicht: (2025)
von: Neitemeier, Pit, et al.
Veröffentlicht: (2025)
Does CLIP Know My Face?
von: Hintersdorf, Dominik, et al.
Veröffentlicht: (2022)
von: Hintersdorf, Dominik, et al.
Veröffentlicht: (2022)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
Making deep neural networks right for the right scientific reasons by interacting with their explanations
von: Schramowski, Patrick, et al.
Veröffentlicht: (2020)
von: Schramowski, Patrick, et al.
Veröffentlicht: (2020)
Representation Matters for Mastering Chess: Improved Feature Representation in AlphaZero Outperforms Switching to Transformers
von: Czech, Johannes, et al.
Veröffentlicht: (2023)
von: Czech, Johannes, et al.
Veröffentlicht: (2023)
The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
von: Mundt, Martin, et al.
Veröffentlicht: (2025)
von: Mundt, Martin, et al.
Veröffentlicht: (2025)
Learning from Less: Guiding Deep Reinforcement Learning with Differentiable Symbolic Planning
von: Ye, Zihan, et al.
Veröffentlicht: (2025)
von: Ye, Zihan, et al.
Veröffentlicht: (2025)
Interpretable end-to-end Neurosymbolic Reinforcement Learning agents
von: Grandien, Nils, et al.
Veröffentlicht: (2024)
von: Grandien, Nils, et al.
Veröffentlicht: (2024)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
von: Friedrich, Felix, et al.
Veröffentlicht: (2024)
von: Friedrich, Felix, et al.
Veröffentlicht: (2024)
SLR: Automated Synthesis for Scalable Logical Reasoning
von: Helff, Lukas, et al.
Veröffentlicht: (2025)
von: Helff, Lukas, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning via Object-Centric Attention
von: Blüml, Jannis, et al.
Veröffentlicht: (2025)
von: Blüml, Jannis, et al.
Veröffentlicht: (2025)
Deep Classifier Mimicry without Data Access
von: Braun, Steven, et al.
Veröffentlicht: (2023)
von: Braun, Steven, et al.
Veröffentlicht: (2023)
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
von: Struppek, Lukas, et al.
Veröffentlicht: (2025)
von: Struppek, Lukas, et al.
Veröffentlicht: (2025)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
von: Karmore, Aryan
Veröffentlicht: (2026)
von: Karmore, Aryan
Veröffentlicht: (2026)
Information Ecosystem Reengineering via Public Sector Knowledge Representation
von: Bagchi, Mayukh
Veröffentlicht: (2025)
von: Bagchi, Mayukh
Veröffentlicht: (2025)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
von: Ostermann, Simon, et al.
Veröffentlicht: (2024)
von: Ostermann, Simon, et al.
Veröffentlicht: (2024)
Causal Explanations Over Time: Articulated Reasoning for Interactive Environments
von: Rödling, Sebastian, et al.
Veröffentlicht: (2025)
von: Rödling, Sebastian, et al.
Veröffentlicht: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
von: Deiseroth, Björn, et al.
Veröffentlicht: (2024) -
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
von: Sztwiertnia, Sebastian, et al.
Veröffentlicht: (2025) -
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
von: Höth, Max Henning, et al.
Veröffentlicht: (2026) -
Core Tokensets for Data-efficient Sequential Training of Transformers
von: Paul, Subarnaduti, et al.
Veröffentlicht: (2024) -
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
von: Helff, Lukas, et al.
Veröffentlicht: (2024)