DreamReader: An Interpretability Toolkit for Text-to-Image Models
Fuente:
arXiv
Saved in:
| Main Authors: | Prakash, Nirmalendu, Oozeer, Narmeen, Lan, Michael, Samkharadze, Luka, Howard, Phillip, Lee, Roy Ka-Wei, Nathawani, Dhruv, Raval, Shivam, Abdullah, Amirali |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Activation Space Interventions Can Be Transferred Between Large Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)
by: Ivanov, Georgi, et al.
Published: (2026)
Interpreting Bias in Large Language Models: A Feature-Based Approach
by: Prakash, Nirmalendu, et al.
Published: (2024)
by: Prakash, Nirmalendu, et al.
Published: (2024)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Bilinear Convolution Decomposition for Causal RL Interpretability
by: Oozeer, Narmeen, et al.
Published: (2024)
by: Oozeer, Narmeen, et al.
Published: (2024)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
by: Harrasse, Abir, et al.
Published: (2025)
by: Harrasse, Abir, et al.
Published: (2025)
Position: Require Frontier AI Labs To Release Small "Analog" Models
by: Upadhyay, Shriyash, et al.
Published: (2025)
by: Upadhyay, Shriyash, et al.
Published: (2025)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
by: Prakash, Nirmalendu, et al.
Published: (2025)
by: Prakash, Nirmalendu, et al.
Published: (2025)
Understanding Refusal in Language Models with Sparse Autoencoders
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
SGHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Singapore
by: Ng, Ri Chi, et al.
Published: (2024)
by: Ng, Ri Chi, et al.
Published: (2024)
Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
by: Gulati, Idhant, et al.
Published: (2026)
by: Gulati, Idhant, et al.
Published: (2026)
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
by: Maltbie, Benjamin, et al.
Published: (2026)
by: Maltbie, Benjamin, et al.
Published: (2026)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
by: Roytburg, Dani, et al.
Published: (2025)
by: Roytburg, Dani, et al.
Published: (2025)
Exploring the Dyson Ring: Parameters, Stability and Helical Orbit
by: Raval, Teerth, et al.
Published: (2024)
by: Raval, Teerth, et al.
Published: (2024)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Interpreting Post colonialism in Ben Okri’s The Famished Road
by: Dhruti Raval
Published: (2019)
by: Dhruti Raval
Published: (2019)
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
by: Yost, Alexandra, et al.
Published: (2025)
by: Yost, Alexandra, et al.
Published: (2025)
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
by: Boxo, Gerard, et al.
Published: (2025)
by: Boxo, Gerard, et al.
Published: (2025)
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
by: Sharma, Aryan, et al.
Published: (2026)
by: Sharma, Aryan, et al.
Published: (2026)
Error Estimation for Adaptive Mesh Refinement in Droplet Simulations
by: Nathawani, Darsh, et al.
Published: (2025)
by: Nathawani, Darsh, et al.
Published: (2025)
A one-dimensional mathematical model for shear-induced droplet formation in co-flowing fluids
by: Nathawani, Darsh, et al.
Published: (2023)
by: Nathawani, Darsh, et al.
Published: (2023)
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
by: Quirke, Philip, et al.
Published: (2025)
by: Quirke, Philip, et al.
Published: (2025)
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation
by: Rosenman, Shachar, et al.
Published: (2023)
by: Rosenman, Shachar, et al.
Published: (2023)
Approximating Human Preferences Using a Multi-Judge Learned System
by: Sprejer, Eitán, et al.
Published: (2025)
by: Sprejer, Eitán, et al.
Published: (2025)
CRISP-NAM: Competing Risks Interpretable Survival Prediction with Neural Additive Models
by: Ramachandram, Dhanesh, et al.
Published: (2025)
by: Ramachandram, Dhanesh, et al.
Published: (2025)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
Caught in the Act: a mechanistic approach to detecting deception
by: Boxo, Gerard, et al.
Published: (2025)
by: Boxo, Gerard, et al.
Published: (2025)
Dreaming of disability‐as‐possibility as a humanistic STEM education futurity
by: Phillip A. Boda
Published: (2024)
by: Phillip A. Boda
Published: (2024)
H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models
by: Dawes, Cutter, et al.
Published: (2026)
by: Dawes, Cutter, et al.
Published: (2026)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
Sycophancy as compositions of Atomic Psychometric Traits
by: Jain, Shreyans, et al.
Published: (2025)
by: Jain, Shreyans, et al.
Published: (2025)
Freud and the Creative Writer: An analysis of Writing, Dreaming and the Interpretation of Dreams
by: Nirjharini Tripathy
Published: (2018)
by: Nirjharini Tripathy
Published: (2018)
Adaptive Length Image Tokenization via Recurrent Allocation
by: Duggal, Shivam, et al.
Published: (2024)
by: Duggal, Shivam, et al.
Published: (2024)
Experimental investigation of laser texturing on surface roughness and wettability of PAHT CF15 fabricated by fused deposition modeling
by: Shivam Prasad, et al.
Published: (2024)
by: Shivam Prasad, et al.
Published: (2024)
DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization
by: Nam, Jisu, et al.
Published: (2024)
by: Nam, Jisu, et al.
Published: (2024)
MoE Lens -- An Expert Is All You Need
by: Chaudhari, Marmik, et al.
Published: (2026)
by: Chaudhari, Marmik, et al.
Published: (2026)
Similar Items
-
Activation Space Interventions Can Be Transferred Between Large Language Models
by: Oozeer, Narmeen, et al.
Published: (2025) -
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025) -
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
by: Prakash, Nirmalendu, et al.
Published: (2026) -
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026) -
Interpreting Bias in Large Language Models: A Feature-Based Approach
by: Prakash, Nirmalendu, et al.
Published: (2024)