Sparse Autoencoders for Hypothesis Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Movva, Rajiv, Peng, Kenny, Garg, Nikhil, Kleinberg, Jon, Pierson, Emma |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
by: Peng, Kenny, et al.
Published: (2025)
by: Peng, Kenny, et al.
Published: (2025)
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023)
by: Movva, Rajiv, et al.
Published: (2023)
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
by: Garg, Nikhil, et al.
Published: (2026)
by: Garg, Nikhil, et al.
Published: (2026)
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
by: Movva, Rajiv, et al.
Published: (2025)
by: Movva, Rajiv, et al.
Published: (2025)
Correlated Errors in Large Language Models
by: Kim, Elliot, et al.
Published: (2025)
by: Kim, Elliot, et al.
Published: (2025)
A No Free Lunch Theorem for Human-AI Collaboration
by: Peng, Kenny, et al.
Published: (2024)
by: Peng, Kenny, et al.
Published: (2024)
Generative AI in Medicine
by: Shanmugam, Divya, et al.
Published: (2024)
by: Shanmugam, Divya, et al.
Published: (2024)
Annotation alignment: Comparing LLM and human annotations of conversational safety
by: Movva, Rajiv, et al.
Published: (2024)
by: Movva, Rajiv, et al.
Published: (2024)
Hypothesis Generation with Large Language Models
by: Zhou, Yangqiaoyu, et al.
Published: (2024)
by: Zhou, Yangqiaoyu, et al.
Published: (2024)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2025)
by: Liu, Haokun, et al.
Published: (2025)
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2024)
by: Liu, Haokun, et al.
Published: (2024)
Learning Disease Progression Models That Capture Health Disparities
by: Chiang, Erica, et al.
Published: (2024)
by: Chiang, Erica, et al.
Published: (2024)
Language Generation in the Limit
by: Kleinberg, Jon, et al.
Published: (2024)
by: Kleinberg, Jon, et al.
Published: (2024)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Evaluating the World Model Implicit in a Generative Model
by: Vafa, Keyon, et al.
Published: (2024)
by: Vafa, Keyon, et al.
Published: (2024)
Modeling the Economic Impacts of AI Openness Regulation
by: Qiu, Tori, et al.
Published: (2025)
by: Qiu, Tori, et al.
Published: (2025)
The Backfiring Effect of Weak AI Safety Regulation
by: Laufer, Benjamin, et al.
Published: (2025)
by: Laufer, Benjamin, et al.
Published: (2025)
The Lock-in Hypothesis: Stagnation by Algorithm
by: Qiu, Tianyi Alex, et al.
Published: (2025)
by: Qiu, Tianyi Alex, et al.
Published: (2025)
The Ontological Dissonance Hypothesis: AI-Triggered Delusional Ideation as Folie a Deux Technologique
by: Lipinska, Izabela, et al.
Published: (2025)
by: Lipinska, Izabela, et al.
Published: (2025)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
by: Xiong, Guangzhi, et al.
Published: (2025)
by: Xiong, Guangzhi, et al.
Published: (2025)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
ASCenD-BDS: Adaptable, Stochastic and Context-aware framework for Detection of Bias, Discrimination and Stereotyping
by: Bahl, Rajiv, et al.
Published: (2025)
by: Bahl, Rajiv, et al.
Published: (2025)
A Bayesian Spatial Model to Correct Under-Reporting in Urban Crowdsourcing
by: Agostini, Gabriel, et al.
Published: (2023)
by: Agostini, Gabriel, et al.
Published: (2023)
Constrain Alignment with Sparse Autoencoders
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
In your own words: computationally identifying interpretable themes in free-text survey data
by: Wang, Jenny S, et al.
Published: (2026)
by: Wang, Jenny S, et al.
Published: (2026)
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Beyond Instrumental and Substitutive Paradigms: Introducing Machine Culture as an Emergent Phenomenon in Large Language Models
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
Clinical Note Bloat Reduction for Efficient LLM Use
by: Cahoon, Jordan L., et al.
Published: (2026)
by: Cahoon, Jordan L., et al.
Published: (2026)
From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models
by: Alsagheer, Dana, et al.
Published: (2025)
by: Alsagheer, Dana, et al.
Published: (2025)
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
by: Shu, Huizhen, et al.
Published: (2025)
by: Shu, Huizhen, et al.
Published: (2025)
Emissions and Performance Trade-off Between Small and Large Language Models
by: Garg, Anandita, et al.
Published: (2025)
by: Garg, Anandita, et al.
Published: (2025)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)
by: Greenwood, Sophie, et al.
Published: (2025)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
by: Liu, Dengcan, et al.
Published: (2025)
by: Liu, Dengcan, et al.
Published: (2025)
Explainability and Certification of AI-Generated Educational Assessments
by: Yaacoub, Antoun, et al.
Published: (2026)
by: Yaacoub, Antoun, et al.
Published: (2026)
Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
by: Laufer, Benjamin, et al.
Published: (2025)
by: Laufer, Benjamin, et al.
Published: (2025)
The Responsible Development of Automated Student Feedback with Generative AI
by: Lindsay, Euan D, et al.
Published: (2023)
by: Lindsay, Euan D, et al.
Published: (2023)
Measuring Human Contribution in AI-Assisted Content Generation
by: Xie, Yueqi, et al.
Published: (2024)
by: Xie, Yueqi, et al.
Published: (2024)
Effect of Gender Fair Job Description on Generative AI Images
by: Böckling, Finn, et al.
Published: (2025)
by: Böckling, Finn, et al.
Published: (2025)
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
by: Zhou, Lexin, et al.
Published: (2025)
by: Zhou, Lexin, et al.
Published: (2025)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
by: Zugecova, Aneta, et al.
Published: (2024)
by: Zugecova, Aneta, et al.
Published: (2024)
Similar Items
-
Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts
by: Peng, Kenny, et al.
Published: (2025) -
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers
by: Movva, Rajiv, et al.
Published: (2023) -
How Many Features Can a Language Model Store Under the Linear Representation Hypothesis?
by: Garg, Nikhil, et al.
Published: (2026) -
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
by: Movva, Rajiv, et al.
Published: (2025) -
Correlated Errors in Large Language Models
by: Kim, Elliot, et al.
Published: (2025)