LatentQA: Teaching LLMs to Decode Activations Into Natural Language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Alexander, Chen, Lijie, Steinhardt, Jacob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Describing Differences in Image Sets with Natural Language
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
von: Jones, Erik, et al.
Veröffentlicht: (2025)
von: Jones, Erik, et al.
Veröffentlicht: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
SyllabusQA: A Course Logistics Question Answering Dataset
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
On The Role of Reasoning in the Identification of Subtle Stereotypes in Natural Language
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
von: Tian, Jacob-Junqi, et al.
Veröffentlicht: (2023)
Interpreting Latent Student Knowledge Representations in Programming Assignments
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
von: Fernandez, Nigel, et al.
Veröffentlicht: (2024)
Causality for Natural Language Processing
von: Jin, Zhijing
Veröffentlicht: (2025)
von: Jin, Zhijing
Veröffentlicht: (2025)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
von: Hashmat, Abdullah, et al.
Veröffentlicht: (2025)
Mass-Producing Failures of Multimodal Systems with Language Models
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026)
von: Luo, Grace, et al.
Veröffentlicht: (2026)
Interpretable Recognition of Cognitive Distortions in Natural Language Texts
von: Kolonin, Anton, et al.
Veröffentlicht: (2025)
von: Kolonin, Anton, et al.
Veröffentlicht: (2025)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
von: Anzenberg, Eitan, et al.
Veröffentlicht: (2025)
von: Anzenberg, Eitan, et al.
Veröffentlicht: (2025)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
von: de Wynter, Adrian, et al.
Veröffentlicht: (2024)
von: de Wynter, Adrian, et al.
Veröffentlicht: (2024)
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
von: Dumitran, Adrian-Marius, et al.
Veröffentlicht: (2025)
von: Dumitran, Adrian-Marius, et al.
Veröffentlicht: (2025)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
von: Noels, Sander, et al.
Veröffentlicht: (2025)
von: Noels, Sander, et al.
Veröffentlicht: (2025)
Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2024)
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2024)
From Imitation to Introspection: Probing Self-Consciousness in Language Models
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
CALM: Culturally Self-Aware Language Models
von: Shen, Lingzhi, et al.
Veröffentlicht: (2026)
von: Shen, Lingzhi, et al.
Veröffentlicht: (2026)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
von: Contestabile, Matilde, et al.
Veröffentlicht: (2025)
von: Contestabile, Matilde, et al.
Veröffentlicht: (2025)
FoundationalASSIST: An Educational Dataset for Foundational Knowledge Tracing and Pedagogical Grounding of LLMs
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
von: Hartmann, David, et al.
Veröffentlicht: (2026)
von: Hartmann, David, et al.
Veröffentlicht: (2026)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
von: Xian, Ruicheng, et al.
Veröffentlicht: (2025)
von: Xian, Ruicheng, et al.
Veröffentlicht: (2025)
Wikipedia in the Era of LLMs: Evolution and Risks
von: Huang, Siming, et al.
Veröffentlicht: (2025)
von: Huang, Siming, et al.
Veröffentlicht: (2025)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
von: Wei, Shou'ang, et al.
Veröffentlicht: (2025)
von: Wei, Shou'ang, et al.
Veröffentlicht: (2025)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Describing Differences in Image Sets with Natural Language
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023) -
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024) -
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
von: Jones, Erik, et al.
Veröffentlicht: (2025) -
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025) -
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)