Discovering Forbidden Topics in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Rager, Can, Wendler, Chris, Gandikota, Rohit, Bau, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Erasing Conceptual Knowledge from Language Models
by: Gandikota, Rohit, et al.
Published: (2024)
by: Gandikota, Rohit, et al.
Published: (2024)
In-Context Algebra
by: Todd, Eric, et al.
Published: (2025)
by: Todd, Eric, et al.
Published: (2025)
DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
by: Kim, Wonyoung, et al.
Published: (2025)
by: Kim, Wonyoung, et al.
Published: (2025)
Can Language Models Discover Scaling Laws?
by: Lin, Haowei, et al.
Published: (2025)
by: Lin, Haowei, et al.
Published: (2025)
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Discovering Latent Knowledge in Language Models Without Supervision
by: Burns, Collin, et al.
Published: (2022)
by: Burns, Collin, et al.
Published: (2022)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
by: Ding, Jiale, et al.
Published: (2025)
by: Ding, Jiale, et al.
Published: (2025)
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning
by: Xu, Yi, et al.
Published: (2026)
by: Xu, Yi, et al.
Published: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Exploring Anti-Aging Literature via ConvexTopics and Large Language Models
by: Yeganova, Lana E., et al.
Published: (2026)
by: Yeganova, Lana E., et al.
Published: (2026)
Finding Culture-Sensitive Neurons in Vision-Language Models
by: Zhao, Xiutian, et al.
Published: (2025)
by: Zhao, Xiutian, et al.
Published: (2025)
Industry-Aligned Granular Topic Modeling
by: Moon, Sae Young, et al.
Published: (2026)
by: Moon, Sae Young, et al.
Published: (2026)
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
by: Li, Kenneth, et al.
Published: (2022)
by: Li, Kenneth, et al.
Published: (2022)
TopicTag: Automatic Annotation of NMF Topic Models Using Chain of Thought and Prompt Tuning with LLMs
by: Wanna, Selma, et al.
Published: (2024)
by: Wanna, Selma, et al.
Published: (2024)
TopicProphet: Prophesies on Temporal Topic Trends and Stocks
by: Kim, Olivia
Published: (2025)
by: Kim, Olivia
Published: (2025)
Forbidden Facts: An Investigation of Competing Objectives in Llama-2
by: Wang, Tony T., et al.
Published: (2023)
by: Wang, Tony T., et al.
Published: (2023)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
by: Zhang, Jinbin, et al.
Published: (2025)
by: Zhang, Jinbin, et al.
Published: (2025)
LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering
by: Wong, Sing Hieng, et al.
Published: (2026)
by: Wong, Sing Hieng, et al.
Published: (2026)
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
by: Huang, Saffron, et al.
Published: (2025)
by: Huang, Saffron, et al.
Published: (2025)
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking
by: Acikgoz, Emre Can, et al.
Published: (2024)
by: Acikgoz, Emre Can, et al.
Published: (2024)
Evaluating Negative Sampling Approaches for Neural Topic Models
by: Adhya, Suman, et al.
Published: (2025)
by: Adhya, Suman, et al.
Published: (2025)
S2WTM: Spherical Sliced-Wasserstein Autoencoder for Topic Modeling
by: Adhya, Suman, et al.
Published: (2025)
by: Adhya, Suman, et al.
Published: (2025)
AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
by: Çano, Erion, et al.
Published: (2024)
by: Çano, Erion, et al.
Published: (2024)
Large Language Models Are Overparameterized Text Encoders
by: K, Thennal D, et al.
Published: (2024)
by: K, Thennal D, et al.
Published: (2024)
TopicDiff: A Topic-enriched Diffusion Approach for Multimodal Conversational Emotion Detection
by: Luo, Jiamin, et al.
Published: (2024)
by: Luo, Jiamin, et al.
Published: (2024)
Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation
by: Yoo, YoungJoon, et al.
Published: (2023)
by: Yoo, YoungJoon, et al.
Published: (2023)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
by: Vincent, Sebastian, et al.
Published: (2023)
by: Vincent, Sebastian, et al.
Published: (2023)
Symmetry-Constrained Language-Guided Program Synthesis for Discovering Governing Equations from Noisy and Partial Observations
by: Baig, Mirza Samad Ahmed, et al.
Published: (2026)
by: Baig, Mirza Samad Ahmed, et al.
Published: (2026)
Hidden State Poisoning Attacks against Mamba-based Language Models
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
Generating Multiple-Choice Knowledge Questions with Interpretable Difficulty Estimation using Knowledge Graphs and Large Language Models
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
by: Şakiroğlu, Mehmet Can, et al.
Published: (2026)
Transformer-based Joint Modelling for Automatic Essay Scoring and Off-Topic Detection
by: Das, Sourya Dipta, et al.
Published: (2024)
by: Das, Sourya Dipta, et al.
Published: (2024)
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
by: Smith, Matthew L., et al.
Published: (2026)
by: Smith, Matthew L., et al.
Published: (2026)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026)
by: Huang, Xinting, et al.
Published: (2026)
Large Language Model Unlearning via Embedding-Corrupted Prompts
by: Liu, Chris Yuhao, et al.
Published: (2024)
by: Liu, Chris Yuhao, et al.
Published: (2024)
Don't Shoot The Breeze: Topic Continuity Model Using Nonlinear Naive Bayes With Attention
by: Pi, Shu-Ting, et al.
Published: (2026)
by: Pi, Shu-Ting, et al.
Published: (2026)
Similar Items
-
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024) -
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024) -
Erasing Conceptual Knowledge from Language Models
by: Gandikota, Rohit, et al.
Published: (2024) -
In-Context Algebra
by: Todd, Eric, et al.
Published: (2025) -
DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
by: Kim, Wonyoung, et al.
Published: (2025)