Saved in:
| Main Authors: | Joseph, Sebastian Antony, Husain, Syed Murtaza, Offner, Stella S. R., Juneau, Stéphanie, Torrey, Paul, Bolton, Adam S., Farias, Juan P., Gaffney, Niall, Durrett, Greg, Li, Junyi Jessy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.20538 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using Natural Language Explanations to Rescale Human Judgments
by: Wadhwa, Manya, et al.
Published: (2023)
by: Wadhwa, Manya, et al.
Published: (2023)
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024)
by: Wadhwa, Manya, et al.
Published: (2024)
VESTA: Visual Exploration with Statistical Tool Agents
by: Rudman, William, et al.
Published: (2026)
by: Rudman, William, et al.
Published: (2026)
Which questions should I answer? Salience Prediction of Inquisitive Questions
by: Wu, Yating, et al.
Published: (2024)
by: Wu, Yating, et al.
Published: (2024)
CREATE: Testing LLMs for Associative Creativity
by: Wadhwa, Manya, et al.
Published: (2026)
by: Wadhwa, Manya, et al.
Published: (2026)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
by: Xin, Yutong, et al.
Published: (2026)
by: Xin, Yutong, et al.
Published: (2026)
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
by: Namuduri, Ramya, et al.
Published: (2025)
by: Namuduri, Ramya, et al.
Published: (2025)
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing
by: Wei, Jiayi, et al.
Published: (2023)
by: Wei, Jiayi, et al.
Published: (2023)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
by: Liu, Zeyu Leo, et al.
Published: (2024)
by: Liu, Zeyu Leo, et al.
Published: (2024)
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
by: Khatry, Anirudh, et al.
Published: (2025)
by: Khatry, Anirudh, et al.
Published: (2025)
AstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
by: Gunjal, Anisha, et al.
Published: (2024)
by: Gunjal, Anisha, et al.
Published: (2024)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
by: Shi, Jinghang, et al.
Published: (2025)
by: Shi, Jinghang, et al.
Published: (2025)
Computational advances and challenges in simulations of turbulence and star formation
by: Federrath, Christoph, et al.
Published: (2025)
by: Federrath, Christoph, et al.
Published: (2025)
AstroPT: Scaling Large Observation Models for Astronomy
by: Smith, Michael J., et al.
Published: (2024)
by: Smith, Michael J., et al.
Published: (2024)
AstroMLab 1: Who Wins Astronomy Jeopardy!?
by: Ting, Yuan-Sen, et al.
Published: (2024)
by: Ting, Yuan-Sen, et al.
Published: (2024)
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
by: Ai, Kuangshi, et al.
Published: (2026)
by: Ai, Kuangshi, et al.
Published: (2026)
Astro-COLIBRI: Empowering Citizen Scientists in Time Domain Astronomy
by: Schüssler, Fabian, et al.
Published: (2024)
by: Schüssler, Fabian, et al.
Published: (2024)
OntoPortal-Astro, a Semantic Artefact Catalogue for Astronomy
by: Cecconi, Baptiste, et al.
Published: (2025)
by: Cecconi, Baptiste, et al.
Published: (2025)
Stellar populations in STARFORGE II: Comparison with observations
by: Farias, Juan P., et al.
Published: (2025)
by: Farias, Juan P., et al.
Published: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
by: Ding, Wenxuan, et al.
Published: (2026)
by: Ding, Wenxuan, et al.
Published: (2026)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
by: Lake, Thom, et al.
Published: (2024)
by: Lake, Thom, et al.
Published: (2024)
LoFiT: Localized Fine-tuning on LLM Representations
by: Yin, Fangcong, et al.
Published: (2024)
by: Yin, Fangcong, et al.
Published: (2024)
CLEVER: A Curated Benchmark for Formally Verified Code Generation
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
Multimodal QUD: Inquisitive Questions from Scientific Figures
by: Wu, Yating, et al.
Published: (2026)
by: Wu, Yating, et al.
Published: (2026)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
by: Hu, Zichao, et al.
Published: (2024)
by: Hu, Zichao, et al.
Published: (2024)
LEGOBench: Scientific Leaderboard Generation Benchmark
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
by: Liu, Zeyu Leo, et al.
Published: (2025)
by: Liu, Zeyu Leo, et al.
Published: (2025)
X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs
by: Rodriguez, Juan Diego, et al.
Published: (2023)
by: Rodriguez, Juan Diego, et al.
Published: (2023)
SuperBench: A Super-Resolution Benchmark Dataset for Scientific Machine Learning
by: Ren, Pu, et al.
Published: (2023)
by: Ren, Pu, et al.
Published: (2023)
AstroSpy: On detecting Fake Images in Astronomy via Joint Image-Spectral Representations
by: Alam, Mohammed Talha, et al.
Published: (2024)
by: Alam, Mohammed Talha, et al.
Published: (2024)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
by: Tripathi, Tuhina, et al.
Published: (2025)
by: Tripathi, Tuhina, et al.
Published: (2025)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
by: Sriram, Aniruddh, et al.
Published: (2024)
by: Sriram, Aniruddh, et al.
Published: (2024)
Similar Items
-
Using Natural Language Explanations to Rescale Human Judgments
by: Wadhwa, Manya, et al.
Published: (2023) -
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024) -
VESTA: Visual Exploration with Statistical Tool Agents
by: Rudman, William, et al.
Published: (2026) -
Which questions should I answer? Salience Prediction of Inquisitive Questions
by: Wu, Yating, et al.
Published: (2024) -
CREATE: Testing LLMs for Associative Creativity
by: Wadhwa, Manya, et al.
Published: (2026)