Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jones, Erik, Patrawala, Arjun, Steinhardt, Jacob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mass-Producing Failures of Multimodal Systems with Language Models
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Approaching Human-Level Forecasting with Language Models
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
von: Lucchetti, Francesca, et al.
Veröffentlicht: (2024)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
Observable Propagation: Uncovering Feature Vectors in Transformers
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2023)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2023)
I Have No Mouth, and I Must Rhyme: Uncovering Internal Phonetic Representations in LLaMA 3.2
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2025)
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2025)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
von: Trauger, Jacob, et al.
Veröffentlicht: (2025)
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026)
von: Luo, Grace, et al.
Veröffentlicht: (2026)
Eliciting Language Model Behaviors with Investigator Agents
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
von: Jones, Graham M., et al.
Veröffentlicht: (2024)
von: Jones, Graham M., et al.
Veröffentlicht: (2024)
Interpreting and Mitigating Unwanted Uncertainty in LLMs
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
Data Science with LLMs and Interpretable Models
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
Transcoders Find Interpretable LLM Feature Circuits
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
von: Dunefsky, Jacob, et al.
Veröffentlicht: (2024)
Uncovering Latent Human Wellbeing in Language Model Embeddings
von: Freire, Pedro, et al.
Veröffentlicht: (2024)
von: Freire, Pedro, et al.
Veröffentlicht: (2024)
How Likely Do LLMs with CoT Mimic Human Reasoning?
von: Bao, Guangsheng, et al.
Veröffentlicht: (2024)
von: Bao, Guangsheng, et al.
Veröffentlicht: (2024)
Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2022)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2022)
On the Thinking-Language Modeling Gap in Large Language Models
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
(How) Do Language Models Track State?
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Uncertainty as Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering
von: Bakman, Yavuz, et al.
Veröffentlicht: (2025)
von: Bakman, Yavuz, et al.
Veröffentlicht: (2025)
Evaluating Computational Accuracy of Large Language Models in Numerical Reasoning Tasks for Healthcare Applications
von: Malghan, Arjun R.
Veröffentlicht: (2025)
von: Malghan, Arjun R.
Veröffentlicht: (2025)
MEMIT-Merge: Addressing MEMIT's Key-Value Conflicts in Same-Subject Batch Editing for LLMs
von: Dong, Zilu, et al.
Veröffentlicht: (2025)
von: Dong, Zilu, et al.
Veröffentlicht: (2025)
CrossTrafficLLM: A Human-Centric Framework for Interpretable Traffic Intelligence via Large Language Model
von: Du, Zeming, et al.
Veröffentlicht: (2025)
von: Du, Zeming, et al.
Veröffentlicht: (2025)
Describing Differences in Image Sets with Natural Language
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
Interpreting the Effects of Quantization on LLMs
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
von: Singh, Manpreet, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mass-Producing Failures of Multimodal Systems with Language Models
von: Tong, Shengbang, et al.
Veröffentlicht: (2023) -
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023) -
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024) -
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
von: Pan, Alexander, et al.
Veröffentlicht: (2024) -
Monitoring Latent World States in Language Models with Propositional Probes
von: Feng, Jiahai, et al.
Veröffentlicht: (2024)