Evaluating Commercial AI Chatbots as News Intermediaries
Fuente:
arXiv
Saved in:
| Main Authors: | Suzgun, Mirac, Shen, Emily, Bianchi, Federico, Spangher, Alexander, Icard, Thomas, Ho, Daniel E., Jurafsky, Dan, Zou, James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
by: Suzgun, Mirac, et al.
Published: (2025)
by: Suzgun, Mirac, et al.
Published: (2025)
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
by: Bianchi, Federico, et al.
Published: (2023)
by: Bianchi, Federico, et al.
Published: (2023)
A Benchmark for Learning to Translate a New Language from One Grammar Book
by: Tanzer, Garrett, et al.
Published: (2023)
by: Tanzer, Garrett, et al.
Published: (2023)
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
by: Dahl, Matthew, et al.
Published: (2024)
by: Dahl, Matthew, et al.
Published: (2024)
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
Cost-of-Pass: An Economic Framework for Evaluating Language Models
by: Erol, Mehmet Hamza, et al.
Published: (2025)
by: Erol, Mehmet Hamza, et al.
Published: (2025)
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
by: Magesh, Varun, et al.
Published: (2024)
by: Magesh, Varun, et al.
Published: (2024)
AI for Scaling Legal Reform: Mapping and Redacting Racial Covenants in Santa Clara County
by: Surani, Faiz, et al.
Published: (2025)
by: Surani, Faiz, et al.
Published: (2025)
Do Language Models Know When They're Hallucinating References?
by: Agrawal, Ayush, et al.
Published: (2023)
by: Agrawal, Ayush, et al.
Published: (2023)
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
NewsEdits 2.0: Learning the Intentions Behind Updating News
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
Explaining Mixtures of Sources in News Articles
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models
by: Zhang, Christine, et al.
Published: (2026)
by: Zhang, Christine, et al.
Published: (2026)
A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models
by: de la Fuente, Antón, et al.
Published: (2024)
by: de la Fuente, Antón, et al.
Published: (2024)
DiscoSum: Discourse-aware News Summarization
by: Spangher, Alexander, et al.
Published: (2025)
by: Spangher, Alexander, et al.
Published: (2025)
Dialect prejudice predicts AI decisions about people's character, employability, and criminality
by: Hofmann, Valentin, et al.
Published: (2024)
by: Hofmann, Valentin, et al.
Published: (2024)
PatentEdits: Framing Patent Novelty as Textual Entailment
by: Lee, Ryan, et al.
Published: (2024)
by: Lee, Ryan, et al.
Published: (2024)
SumTablets: A Transliteration Dataset of Sumerian Tablets
by: Simmons, Cole, et al.
Published: (2026)
by: Simmons, Cole, et al.
Published: (2026)
HumT DumT: Measuring and controlling human-like language in LLMs
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
by: Bianchi, Federico, et al.
Published: (2025)
by: Bianchi, Federico, et al.
Published: (2025)
Othering and low status framing of immigrant cuisines in US restaurant reviews and large language models
by: Luo, Yiwei, et al.
Published: (2023)
by: Luo, Yiwei, et al.
Published: (2023)
NewsHomepages: Homepage Layouts Capture Information Prioritization Decisions
by: Welsh, Ben, et al.
Published: (2024)
by: Welsh, Ben, et al.
Published: (2024)
Humans overrely on overconfident language models, across languages
by: Rathi, Neil, et al.
Published: (2025)
by: Rathi, Neil, et al.
Published: (2025)
Automated Benchmark Auditing for AI Agents and Large Language Models
by: Wang, Junlin, et al.
Published: (2026)
by: Wang, Junlin, et al.
Published: (2026)
False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models
by: Kallini, Julie, et al.
Published: (2025)
by: Kallini, Julie, et al.
Published: (2025)
NewsInterview: a Dataset and a Playground to Evaluate LLMs' Ground Gap via Informational Interviews
by: Spangher, Alexander, et al.
Published: (2024)
by: Spangher, Alexander, et al.
Published: (2024)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
by: Zhou, Kaitlyn, et al.
Published: (2026)
by: Zhou, Kaitlyn, et al.
Published: (2026)
How Causal Abstraction Underpins Computational Explanation
by: Geiger, Atticus, et al.
Published: (2025)
by: Geiger, Atticus, et al.
Published: (2025)
Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
by: Ògúnrèmí, Tolúlopé, et al.
Published: (2025)
by: Ògúnrèmí, Tolúlopé, et al.
Published: (2025)
Data Checklist: On Unit-Testing Datasets with Usable Information
by: Zhang, Heidi C., et al.
Published: (2024)
by: Zhang, Heidi C., et al.
Published: (2024)
Advancing Academic Chatbots: Evaluation of Non Traditional Outputs
by: Favero, Nicole, et al.
Published: (2025)
by: Favero, Nicole, et al.
Published: (2025)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
by: Arora, Aryaman, et al.
Published: (2024)
by: Arora, Aryaman, et al.
Published: (2024)
From Noise to Signal: When Outliers Seed New Topics
by: Zve, Evangelia, et al.
Published: (2026)
by: Zve, Evangelia, et al.
Published: (2026)
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
by: Cheng, Myra, et al.
Published: (2024)
by: Cheng, Myra, et al.
Published: (2024)
Beyond Tokens: Concept-Level Training Objectives for LLMs
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
by: Shani, Chen, et al.
Published: (2026)
by: Shani, Chen, et al.
Published: (2026)
Interleaving Logic and Counting
by: van Benthem, Johan, et al.
Published: (2025)
by: van Benthem, Johan, et al.
Published: (2025)
Similar Items
-
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024) -
Dynamic Cheatsheet: Test-Time Learning with Adaptive Memory
by: Suzgun, Mirac, et al.
Published: (2025) -
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
by: Bianchi, Federico, et al.
Published: (2023) -
A Benchmark for Learning to Translate a New Language from One Grammar Book
by: Tanzer, Garrett, et al.
Published: (2023) -
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
by: Dahl, Matthew, et al.
Published: (2024)