EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cappello, Franck, Madireddy, Sandeep, Underwood, Robert, Getty, Neil, Chia, Nicholas Lee-Ping, Ramachandra, Nesar, Nguyen, Josh, Keceli, Murat, Mallick, Tanwi, Li, Zilinghan, Ngom, Marieme, Zhang, Chenhui, Yanguas-Gil, Angel, Antoniuk, Evan, Kailkhura, Bhavya, Tian, Minyang, Du, Yufeng, Ting, Yuan-Sen, Wells, Azton, Nicolae, Bogdan, Maurya, Avinash, Rafique, M. Mustafa, Huerta, Eliu, Li, Bo, Foster, Ian, Stevens, Rick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
von: Maurya, Avinash, et al.
Veröffentlicht: (2025)
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Kernel Model Validation: How To Do It, And Why You Should Care
von: Graziani, Carlo, et al.
Veröffentlicht: (2025)
von: Graziani, Carlo, et al.
Veröffentlicht: (2025)
Targeted Adaptive Design
von: Graziani, Carlo, et al.
Veröffentlicht: (2022)
von: Graziani, Carlo, et al.
Veröffentlicht: (2022)
Teaching LLMs to Speak Spectroscopy
von: Ramachandra, Nesar, et al.
Veröffentlicht: (2025)
von: Ramachandra, Nesar, et al.
Veröffentlicht: (2025)
Multi-modal Foundation Model for Cosmological Simulation Data
von: Xia, Bin, et al.
Veröffentlicht: (2025)
von: Xia, Bin, et al.
Veröffentlicht: (2025)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
von: Gokdemir, Ozan, et al.
Veröffentlicht: (2025)
von: Gokdemir, Ozan, et al.
Veröffentlicht: (2025)
UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-Making
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
Extending $μ$P: Spectral Conditions for Feature Learning Across Optimizers
von: Gupta, Akshita, et al.
Veröffentlicht: (2026)
von: Gupta, Akshita, et al.
Veröffentlicht: (2026)
Uncovering Physical Drivers of Dark Matter Halo Structures with Auxiliary-Variable-Guided Generative Models
von: Ganguli, Arkaprabha, et al.
Veröffentlicht: (2026)
von: Ganguli, Arkaprabha, et al.
Veröffentlicht: (2026)
AstroMLab 1: Who Wins Astronomy Jeopardy!?
von: Ting, Yuan-Sen, et al.
Veröffentlicht: (2024)
von: Ting, Yuan-Sen, et al.
Veröffentlicht: (2024)
Active Learning Enables Extrapolation in Molecular Generative Models
von: Antoniuk, Evan R., et al.
Veröffentlicht: (2025)
von: Antoniuk, Evan R., et al.
Veröffentlicht: (2025)
MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2025)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2025)
Wavelet-Inspired Multiscale Graph Convolutional Recurrent Network for Traffic Forecasting
von: Qian, Qipeng, et al.
Veröffentlicht: (2024)
von: Qian, Qipeng, et al.
Veröffentlicht: (2024)
No Test Cases, No Problem: Distillation-Driven Code Generation for Scientific Workflows
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2026)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2026)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
von: Du, Yufeng, et al.
Veröffentlicht: (2025)
From Atomistic Models to Machine Learning: Predictive Design of Nanocarbons under Extreme Conditions
von: Yan, Xiaoli, et al.
Veröffentlicht: (2026)
von: Yan, Xiaoli, et al.
Veröffentlicht: (2026)
Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
von: Yang, Chih-Hsuan, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Hsuan, et al.
Veröffentlicht: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Wang, Zhengyang, et al.
Veröffentlicht: (2025)
Benchmarking AI-evolved cosmological structure formation
von: Dong, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Dong, Xiaofeng, et al.
Veröffentlicht: (2025)
Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets
von: Ganguli, Arkaprabha, et al.
Veröffentlicht: (2025)
von: Ganguli, Arkaprabha, et al.
Veröffentlicht: (2025)
LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals
von: Yeh, Samuel, et al.
Veröffentlicht: (2025)
von: Yeh, Samuel, et al.
Veröffentlicht: (2025)
AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model
von: de Haan, Tijmen, et al.
Veröffentlicht: (2024)
von: de Haan, Tijmen, et al.
Veröffentlicht: (2024)
Reducing Model Error Using Optimised Galaxy Selection: Weak Lensing Cluster Mass Estimation
von: Rau, Markus Michael, et al.
Veröffentlicht: (2024)
von: Rau, Markus Michael, et al.
Veröffentlicht: (2024)
Toward Reliable, Safe, and Secure LLMs for Scientific Applications
von: Chaturvedi, Saket Sanjeev, et al.
Veröffentlicht: (2026)
von: Chaturvedi, Saket Sanjeev, et al.
Veröffentlicht: (2026)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
FedSZ: Leveraging Error-Bounded Lossy Compression for Federated Learning Communications
von: Wilkins, Grant, et al.
Veröffentlicht: (2023)
von: Wilkins, Grant, et al.
Veröffentlicht: (2023)
AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model
von: de Haan, Tijmen, et al.
Veröffentlicht: (2025)
von: de Haan, Tijmen, et al.
Veröffentlicht: (2025)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
FedCluster: Boosting the Convergence of Federated Learning via Cluster-Cycling
von: Chen, Cheng, et al.
Veröffentlicht: (2020)
von: Chen, Cheng, et al.
Veröffentlicht: (2020)
Improving Robustness In Sparse Autoencoders via Masked Regularization
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
von: Narayanaswamy, Vivek, et al.
Veröffentlicht: (2026)
Certifiably-Robust Federated Adversarial Learning via Randomized Smoothing
von: Chen, Cheng, et al.
Veröffentlicht: (2021)
von: Chen, Cheng, et al.
Veröffentlicht: (2021)
Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
von: Yang, Hongru, et al.
Veröffentlicht: (2024)
von: Yang, Hongru, et al.
Veröffentlicht: (2024)
End-to-End Mesh Optimization of a Hybrid Deep Learning Black-Box PDE Solver
von: Ma, Shaocong, et al.
Veröffentlicht: (2024)
von: Ma, Shaocong, et al.
Veröffentlicht: (2024)
Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints
von: Balaji, Adarsha, et al.
Veröffentlicht: (2025)
von: Balaji, Adarsha, et al.
Veröffentlicht: (2025)
Multi-task Modeling for Engineering Applications with Sparse Data
von: Comlek, Yigitcan, et al.
Veröffentlicht: (2026)
von: Comlek, Yigitcan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024) -
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026) -
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
von: Maurya, Avinash, et al.
Veröffentlicht: (2025) -
Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
von: Maurya, Avinash, et al.
Veröffentlicht: (2024) -
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)