Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bajpai, Prasoon, Chatterjee, Niladri, Dutta, Subhabrata, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$\texttt{LM}^\texttt{2}$: A Simple Society of Language Models Solves Complex Reasoning
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
Mechanistic Behavior Editing of Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024)
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
von: Hengle, Amey, et al.
Veröffentlicht: (2025)
von: Hengle, Amey, et al.
Veröffentlicht: (2025)
Multilingual Test-Time Scaling via Initial Thought Transfer
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2025)
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2025)
Small Language Models Fine-tuned to Coordinate Larger Language Models improve Complex Reasoning
von: Juneja, Gurusha, et al.
Veröffentlicht: (2023)
von: Juneja, Gurusha, et al.
Veröffentlicht: (2023)
Information Anxiety in Large Language Models
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2024)
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2024)
Can LLMs Automate Fact-Checking Article Writing?
von: Sahnan, Dhruv, et al.
Veröffentlicht: (2025)
von: Sahnan, Dhruv, et al.
Veröffentlicht: (2025)
HIDE and Seek: Detecting Hallucinations in Language Models via Decoupled Representations
von: Chatterjee, Anwoy, et al.
Veröffentlicht: (2025)
von: Chatterjee, Anwoy, et al.
Veröffentlicht: (2025)
Language Models can Exploit Cross-Task In-context Learning for Data-Scarce Novel Tasks
von: Chatterjee, Anwoy, et al.
Veröffentlicht: (2024)
von: Chatterjee, Anwoy, et al.
Veröffentlicht: (2024)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs
von: Hengle, Amey, et al.
Veröffentlicht: (2025)
von: Hengle, Amey, et al.
Veröffentlicht: (2025)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
von: Chatrath, Veronica, et al.
Veröffentlicht: (2024)
von: Chatrath, Veronica, et al.
Veröffentlicht: (2024)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
von: Piot, Paloma, et al.
Veröffentlicht: (2025)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
von: Lunardi, Riccardo, et al.
Veröffentlicht: (2025)
von: Lunardi, Riccardo, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
von: Pres, Itamar, et al.
Veröffentlicht: (2024)
von: Pres, Itamar, et al.
Veröffentlicht: (2024)
Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
von: Zhou, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2025)
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
von: Srivatsa, KV Aditya, et al.
Veröffentlicht: (2025)
von: Srivatsa, KV Aditya, et al.
Veröffentlicht: (2025)
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
von: Doostmohammadi, Ehsan, et al.
Veröffentlicht: (2024)
von: Doostmohammadi, Ehsan, et al.
Veröffentlicht: (2024)
Generating Hierarchical JSON Representations of Scientific Sentences Using LLMs
von: Nimmagadda, Satya Sri Rajiteswari, et al.
Veröffentlicht: (2026)
von: Nimmagadda, Satya Sri Rajiteswari, et al.
Veröffentlicht: (2026)
HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue
von: Iyer, Laya, et al.
Veröffentlicht: (2026)
von: Iyer, Laya, et al.
Veröffentlicht: (2026)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
von: Ma, Huanhuan, et al.
Veröffentlicht: (2025)
von: Ma, Huanhuan, et al.
Veröffentlicht: (2025)
Characterizing and Evaluating the Reliability of LLMs against Jailbreak Attacks
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
von: Chen, Kexin, et al.
Veröffentlicht: (2024)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
von: Saha, Sougata, et al.
Veröffentlicht: (2025)
von: Saha, Sougata, et al.
Veröffentlicht: (2025)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
ExpressivityBench: Can LLMs Communicate Implicitly?
von: Tint, Joshua, et al.
Veröffentlicht: (2024)
von: Tint, Joshua, et al.
Veröffentlicht: (2024)
Semantic Consistency for Assuring Reliability of Large Language Models
von: Raj, Harsh, et al.
Veröffentlicht: (2023)
von: Raj, Harsh, et al.
Veröffentlicht: (2023)
LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2025)
von: Guo, Pei-Fu, et al.
Veröffentlicht: (2025)
DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
von: Shu, Fan, et al.
Veröffentlicht: (2026)
von: Shu, Fan, et al.
Veröffentlicht: (2026)
Can GNN be Good Adapter for LLMs?
von: Huang, Xuanwen, et al.
Veröffentlicht: (2024)
von: Huang, Xuanwen, et al.
Veröffentlicht: (2024)
Can LLMs Ask Good Questions?
von: Zhang, Yueheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yueheng, et al.
Veröffentlicht: (2025)
Can LLMs Explain Themselves Counterfactually?
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2025)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2025)
Can LLMs Capture Human Preferences?
von: Goli, Ali, et al.
Veröffentlicht: (2023)
von: Goli, Ali, et al.
Veröffentlicht: (2023)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2025)
von: Javaji, Shashidhar Reddy, et al.
Veröffentlicht: (2025)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
How Reliable are LLMs for Reasoning on the Re-ranking task?
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2025)
von: Islam, Nafis Tanveer, et al.
Veröffentlicht: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
von: Sundriyal, Megha, et al.
Veröffentlicht: (2023)
von: Sundriyal, Megha, et al.
Veröffentlicht: (2023)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
$\texttt{LM}^\texttt{2}$: A Simple Society of Language Models Solves Complex Reasoning
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024) -
Mechanistic Behavior Editing of Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2024) -
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2024) -
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
von: Hengle, Amey, et al.
Veröffentlicht: (2025) -
Multilingual Test-Time Scaling via Initial Thought Transfer
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2025)