Explaining Datasets in Words: Statistical Models with Natural Language Parameters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Ruiqi, Wang, Heng, Klein, Dan, Steinhardt, Jacob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
von: Feng, Jiahai, et al.
Veröffentlicht: (2023)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Which Attention Heads Matter for In-Context Learning?
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
von: Yin, Kayo, et al.
Veröffentlicht: (2025)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
von: Ye, Yaowen, et al.
Veröffentlicht: (2025)
Approaching Human-Level Forecasting with Language Models
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
Word Embeddings Are Steers for Language Models
von: Han, Chi, et al.
Veröffentlicht: (2023)
von: Han, Chi, et al.
Veröffentlicht: (2023)
Learning a Generative Meta-Model of LLM Activations
von: Luo, Grace, et al.
Veröffentlicht: (2026)
von: Luo, Grace, et al.
Veröffentlicht: (2026)
Eliciting Language Model Behaviors with Investigator Agents
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2025)
Fairness Definitions in Language Models Explained
von: Yin, Zhipeng, et al.
Veröffentlicht: (2024)
von: Yin, Zhipeng, et al.
Veröffentlicht: (2024)
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
von: Bratulić, Jelena, et al.
Veröffentlicht: (2025)
Explaining Large Language Models with gSMILE
von: Dehghani, Zeinab, et al.
Veröffentlicht: (2025)
von: Dehghani, Zeinab, et al.
Veröffentlicht: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
von: Huang, Vincent, et al.
Veröffentlicht: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
Parameter Efficient Fine-tuning via Explained Variance Adaptation
von: Paischer, Fabian, et al.
Veröffentlicht: (2024)
von: Paischer, Fabian, et al.
Veröffentlicht: (2024)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
von: Kim, Minsang, et al.
Veröffentlicht: (2026)
von: Kim, Minsang, et al.
Veröffentlicht: (2026)
Explicit Word Density Estimation for Language Modelling
von: Andonov, Jovan, et al.
Veröffentlicht: (2024)
von: Andonov, Jovan, et al.
Veröffentlicht: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
Learning to Model the World with Language
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
von: Wu, Bohao, et al.
Veröffentlicht: (2025)
von: Wu, Bohao, et al.
Veröffentlicht: (2025)
Explingo: Explaining AI Predictions using Large Language Models
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
von: Zytek, Alexandra, et al.
Veröffentlicht: (2024)
Explaining Large Language Models Decisions Using Shapley Values
von: Mohammadi, Behnam
Veröffentlicht: (2024)
von: Mohammadi, Behnam
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning for Foundation Models
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
Protein Language Models Diverge from Natural Language: Comparative Analysis and Improved Inference
von: Hart, Anna, et al.
Veröffentlicht: (2026)
von: Hart, Anna, et al.
Veröffentlicht: (2026)
Research on Optimization of Natural Language Processing Model Based on Multimodal Deep Learning
von: Sun, Dan, et al.
Veröffentlicht: (2024)
von: Sun, Dan, et al.
Veröffentlicht: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Li, Xiaopeng, et al.
Veröffentlicht: (2024)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
von: Halawi, Danny, et al.
Veröffentlicht: (2024)
World Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word Embeddings
von: Barenholtz, Elan
Veröffentlicht: (2026)
von: Barenholtz, Elan
Veröffentlicht: (2026)
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
Causal Evaluation of Language Models
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
von: Chen, Sirui, et al.
Veröffentlicht: (2024)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Reinterpreting 'the Company a Word Keeps': Towards Explainable and Ontologically Grounded Language Models
von: Saba, Walid S.
Veröffentlicht: (2024)
von: Saba, Walid S.
Veröffentlicht: (2024)
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Yang, Zhuonan, et al.
Veröffentlicht: (2026)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
von: Bonthu, Sridevi, et al.
Veröffentlicht: (2025)
von: Bonthu, Sridevi, et al.
Veröffentlicht: (2025)
Tree Matching Networks for Natural Language Inference: Parameter-Efficient Semantic Understanding via Dependency Parse Trees
von: Lunder, Jason
Veröffentlicht: (2025)
von: Lunder, Jason
Veröffentlicht: (2025)
Ähnliche Einträge
-
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022) -
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025) -
How do Language Models Bind Entities in Context?
von: Feng, Jiahai, et al.
Veröffentlicht: (2023) -
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023) -
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)