What Evidence Do Language Models Find Convincing?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wan, Alexander, Wallace, Eric, Klein, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Predicting Emergent Capabilities by Finetuning
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
von: Noels, Sander, et al.
Veröffentlicht: (2025)
von: Noels, Sander, et al.
Veröffentlicht: (2025)
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
von: Krishna, Kundan, et al.
Veröffentlicht: (2024)
von: Krishna, Kundan, et al.
Veröffentlicht: (2024)
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)
von: Burns, Collin, et al.
Veröffentlicht: (2022)
Function Vectors in Large Language Models
von: Todd, Eric, et al.
Veröffentlicht: (2023)
von: Todd, Eric, et al.
Veröffentlicht: (2023)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhong, Ruiqi, et al.
Veröffentlicht: (2024)
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
von: Klein, Tassilo, et al.
Veröffentlicht: (2024)
von: Klein, Tassilo, et al.
Veröffentlicht: (2024)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
von: Kang, Katie, et al.
Veröffentlicht: (2024)
von: Kang, Katie, et al.
Veröffentlicht: (2024)
Compared to What? Baselines and Metrics for Counterfactual Prompting
von: Yang, Zihao, et al.
Veröffentlicht: (2026)
von: Yang, Zihao, et al.
Veröffentlicht: (2026)
Do Large Language Models Need Intent? Revisiting Response Generation Strategies for Service Assistant
von: Bolshinsky, Inbal, et al.
Veröffentlicht: (2025)
von: Bolshinsky, Inbal, et al.
Veröffentlicht: (2025)
Learning to Model the World with Language
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
von: Lin, Jessy, et al.
Veröffentlicht: (2023)
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
von: Li, Michael, et al.
Veröffentlicht: (2025)
von: Li, Michael, et al.
Veröffentlicht: (2025)
A Single-Layer Model Can Do Language Modeling
von: Wang, Zanmin
Veröffentlicht: (2026)
von: Wang, Zanmin
Veröffentlicht: (2026)
Do Large Language Model Benchmarks Test Reliability?
von: Vendrow, Joshua, et al.
Veröffentlicht: (2025)
von: Vendrow, Joshua, et al.
Veröffentlicht: (2025)
What Do Large Language Models Know? Tacit Knowledge as a Potential Causal-Explanatory Structure
von: Budding, Céline
Veröffentlicht: (2025)
von: Budding, Céline
Veröffentlicht: (2025)
Do Activation Verbalization Methods Convey Privileged Information?
von: Li, Millicent, et al.
Veröffentlicht: (2025)
von: Li, Millicent, et al.
Veröffentlicht: (2025)
What is Wrong with Perplexity for Long-context Language Modeling?
von: Fang, Lizhe, et al.
Veröffentlicht: (2024)
von: Fang, Lizhe, et al.
Veröffentlicht: (2024)
Can Language Models Recognize Convincing Arguments?
von: Rescala, Paula, et al.
Veröffentlicht: (2024)
von: Rescala, Paula, et al.
Veröffentlicht: (2024)
Do LLMs Find Human Answers To Fact-Driven Questions Perplexing? A Case Study on Reddit
von: Seegmiller, Parker, et al.
Veröffentlicht: (2024)
von: Seegmiller, Parker, et al.
Veröffentlicht: (2024)
Unforgettable Generalization in Language Models
von: Zhang, Eric, et al.
Veröffentlicht: (2024)
von: Zhang, Eric, et al.
Veröffentlicht: (2024)
What Will My Model Forget? Forecasting Forgotten Examples in Language Model Refinement
von: Jin, Xisen, et al.
Veröffentlicht: (2024)
von: Jin, Xisen, et al.
Veröffentlicht: (2024)
DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
von: Hu, Xiaolin, et al.
Veröffentlicht: (2024)
von: Hu, Xiaolin, et al.
Veröffentlicht: (2024)
Discursive Circuits: How Do Language Models Understand Discourse Relations?
von: Miao, Yisong, et al.
Veröffentlicht: (2025)
von: Miao, Yisong, et al.
Veröffentlicht: (2025)
What Do Language Models Hear? Probing for Auditory Representations in Language Models
von: Ngo, Jerry, et al.
Veröffentlicht: (2024)
von: Ngo, Jerry, et al.
Veröffentlicht: (2024)
Structural Pruning of Pre-trained Language Models via Neural Architecture Search
von: Klein, Aaron, et al.
Veröffentlicht: (2024)
von: Klein, Aaron, et al.
Veröffentlicht: (2024)
Finding Culture-Sensitive Neurons in Vision-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2025)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2025)
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
von: Bhaskar, Adithya, et al.
Veröffentlicht: (2024)
von: Bhaskar, Adithya, et al.
Veröffentlicht: (2024)
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
von: Min, Sewon, et al.
Veröffentlicht: (2023)
von: Min, Sewon, et al.
Veröffentlicht: (2023)
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
von: Banerjee, Sourav, et al.
Veröffentlicht: (2024)
Comparison of Scoring Rationales Between Large Language Models and Human Raters
von: Hua, Haowei, et al.
Veröffentlicht: (2025)
von: Hua, Haowei, et al.
Veröffentlicht: (2025)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
von: Zverev, Egor, et al.
Veröffentlicht: (2024)
von: Zverev, Egor, et al.
Veröffentlicht: (2024)
Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
von: AlMarri, Saeed, et al.
Veröffentlicht: (2025)
von: AlMarri, Saeed, et al.
Veröffentlicht: (2025)
Protected group bias and stereotypes in Large Language Models
von: Kotek, Hadas, et al.
Veröffentlicht: (2024)
von: Kotek, Hadas, et al.
Veröffentlicht: (2024)
Large Language Model Prompt Datasets: An In-depth Analysis and Insights
von: Zhang, Yuanming, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanming, et al.
Veröffentlicht: (2025)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task
von: Curth, Alicia, et al.
Veröffentlicht: (2026)
von: Curth, Alicia, et al.
Veröffentlicht: (2026)
Fantastic Biases (What are They) and Where to Find Them
von: Barriere, Valentin
Veröffentlicht: (2024)
von: Barriere, Valentin
Veröffentlicht: (2024)
EvidenceRL: Reinforcing Evidence Consistency for Trustworthy Language Models
von: Tamo, J. Ben, et al.
Veröffentlicht: (2026)
von: Tamo, J. Ben, et al.
Veröffentlicht: (2026)
ParaScopes: What do Language Models Activations Encode About Future Text?
von: Pochinkov, Nicky, et al.
Veröffentlicht: (2025)
von: Pochinkov, Nicky, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Predicting Emergent Capabilities by Finetuning
von: Snell, Charlie, et al.
Veröffentlicht: (2024) -
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024) -
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
von: Noels, Sander, et al.
Veröffentlicht: (2025) -
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
von: Krishna, Kundan, et al.
Veröffentlicht: (2024) -
Discovering Latent Knowledge in Language Models Without Supervision
von: Burns, Collin, et al.
Veröffentlicht: (2022)