Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Eger, Steffen, Cao, Yong, D'Souza, Jennifer, Geiger, Andreas, Greisinger, Christian, Gross, Stephanie, Hou, Yufang, Krenn, Brigitte, Lauscher, Anne, Li, Yizhi, Lin, Chenghua, Moosavi, Nafise Sadat, Zhao, Wei, Miller, Tristan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
von: Greisinger, Christian, et al.
Veröffentlicht: (2026)
von: Greisinger, Christian, et al.
Veröffentlicht: (2026)
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
von: Liu, Yiqi, et al.
Veröffentlicht: (2023)
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
von: James, Joseph, et al.
Veröffentlicht: (2026)
von: James, Joseph, et al.
Veröffentlicht: (2026)
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
von: Belouadi, Jonas, et al.
Veröffentlicht: (2023)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2023)
How to Leverage Digit Embeddings to Represent Numbers?
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2024)
von: Sivakumar, Jasivan Alex, et al.
Veröffentlicht: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
MultiHoax: A Dataset of Multi-hop False-Premise Questions
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
von: Mi, Maggie, et al.
Veröffentlicht: (2024)
von: Mi, Maggie, et al.
Veröffentlicht: (2024)
From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
von: Mi, Maggie, et al.
Veröffentlicht: (2025)
von: Mi, Maggie, et al.
Veröffentlicht: (2025)
More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
von: Shafiei, Mohammadamin, et al.
Veröffentlicht: (2025)
Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
von: Pandya, Mugdha, et al.
Veröffentlicht: (2024)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
von: Xue, Huiyin, et al.
Veröffentlicht: (2025)
von: Xue, Huiyin, et al.
Veröffentlicht: (2025)
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection
von: Pastorino, Valeria, et al.
Veröffentlicht: (2024)
von: Pastorino, Valeria, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph
von: Nechakhin, Vladyslav, et al.
Veröffentlicht: (2024)
von: Nechakhin, Vladyslav, et al.
Veröffentlicht: (2024)
No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
von: Permadi, Vynska Amalia, et al.
Veröffentlicht: (2026)
von: Permadi, Vynska Amalia, et al.
Veröffentlicht: (2026)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2025)
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2025)
Faithful Summarisation under Disagreement via Belief-Level Aggregation
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
von: Aghaebe, Favour Yahdii, et al.
Veröffentlicht: (2026)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
von: Saffari, Hamidreza, et al.
Veröffentlicht: (2024)
Exploring Gender Disparities in Automatic Speech Recognition Technology
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
von: ElGhazaly, Hend, et al.
Veröffentlicht: (2025)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
von: Stacey, Joe, et al.
Veröffentlicht: (2026)
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
von: Belouadi, Jonas, et al.
Veröffentlicht: (2024)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2024)
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
von: Kunilovskaya, Maria, et al.
Veröffentlicht: (2026)
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks
von: Chen, Yanran, et al.
Veröffentlicht: (2025)
von: Chen, Yanran, et al.
Veröffentlicht: (2025)
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
von: Leiter, Christoph, et al.
Veröffentlicht: (2024)
USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2022)
LLM-based multi-agent poetry generation in non-cooperative environments
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
BatchGEMBA: Token-Efficient Machine Translation Evaluation with Batched Prompting and Prompt Compression
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
von: Larionov, Daniil, et al.
Veröffentlicht: (2025)
Is there really a Citation Age Bias in NLP?
von: Nguyen, Hoa, et al.
Veröffentlicht: (2024)
von: Nguyen, Hoa, et al.
Veröffentlicht: (2024)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
von: Glockner, Max, et al.
Veröffentlicht: (2024)
von: Glockner, Max, et al.
Veröffentlicht: (2024)
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation
von: Zhang, Ran, et al.
Veröffentlicht: (2023)
von: Zhang, Ran, et al.
Veröffentlicht: (2023)
BMX: Boosting Natural Language Generation Metrics with Explainability
von: Leiter, Christoph, et al.
Veröffentlicht: (2022)
von: Leiter, Christoph, et al.
Veröffentlicht: (2022)
ClaimFlow: Tracing the Evolution of Scientific Claims in NLP
von: Pramanick, Aniket, et al.
Veröffentlicht: (2026)
von: Pramanick, Aniket, et al.
Veröffentlicht: (2026)
Large Language Models as Evaluators for Scientific Synthesis
von: Evans, Julia, et al.
Veröffentlicht: (2024)
von: Evans, Julia, et al.
Veröffentlicht: (2024)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
von: Roy, Subhadeep, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
von: Greisinger, Christian, et al.
Veröffentlicht: (2026) -
LLMs as Narcissistic Evaluators: When Ego Inflates Evaluation Scores
von: Liu, Yiqi, et al.
Veröffentlicht: (2023) -
ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation
von: Wang, Xiao, et al.
Veröffentlicht: (2025) -
RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation
von: James, Joseph, et al.
Veröffentlicht: (2026) -
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
von: Belouadi, Jonas, et al.
Veröffentlicht: (2023)