Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xinge, Zhang, Terry Jingchen, Schölkopf, Bernhard, Jin, Zhijing, Menou, Kristen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Causality can systematically address the monsters under the bench(marks)
von: Leeb, Felix, et al.
Veröffentlicht: (2025)
von: Leeb, Felix, et al.
Veröffentlicht: (2025)
Can Theoretical Physics Research Benefit from Language Agents?
von: Lu, Sirui, et al.
Veröffentlicht: (2025)
von: Lu, Sirui, et al.
Veröffentlicht: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
Causal Responsibility Attribution for Human-AI Collaboration
von: Qi, Yahang, et al.
Veröffentlicht: (2024)
von: Qi, Yahang, et al.
Veröffentlicht: (2024)
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
von: Binkyte, Ruta, et al.
Veröffentlicht: (2025)
von: Binkyte, Ruta, et al.
Veröffentlicht: (2025)
Computational Arbitrage in AI Model Markets
von: Olmedo, Ricardo, et al.
Veröffentlicht: (2026)
von: Olmedo, Ricardo, et al.
Veröffentlicht: (2026)
Decomposing and Measuring Evaluation Awareness
von: Li, Changling, et al.
Veröffentlicht: (2026)
von: Li, Changling, et al.
Veröffentlicht: (2026)
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
von: Zhang, Terry Jingchen, et al.
Veröffentlicht: (2025)
von: Zhang, Terry Jingchen, et al.
Veröffentlicht: (2025)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
Can Large Language Models Infer Causation from Correlation?
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution
von: Binkyte, Ruta, et al.
Veröffentlicht: (2026)
von: Binkyte, Ruta, et al.
Veröffentlicht: (2026)
Out-of-Variable Generalization for Discriminative Models
von: Guo, Siyuan, et al.
Veröffentlicht: (2023)
von: Guo, Siyuan, et al.
Veröffentlicht: (2023)
A Probabilistic Model Behind Self-Supervised Learning
von: Bizeul, Alice, et al.
Veröffentlicht: (2024)
von: Bizeul, Alice, et al.
Veröffentlicht: (2024)
Orthogonal Finetuning Made Scalable
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents
von: Koblischke, Nolan, et al.
Veröffentlicht: (2025)
von: Koblischke, Nolan, et al.
Veröffentlicht: (2025)
Conformal Generative Modeling with Improved Sample Efficiency through Sequential Greedy Filtering
von: Kladny, Klaus-Rudolf, et al.
Veröffentlicht: (2024)
von: Kladny, Klaus-Rudolf, et al.
Veröffentlicht: (2024)
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
von: Cobben, Pepijn, et al.
Veröffentlicht: (2026)
von: Cobben, Pepijn, et al.
Veröffentlicht: (2026)
Implicit Personalization in Language Models: A Systematic Study
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI
von: Huang, Xuanqiang Angelo, et al.
Veröffentlicht: (2026)
von: Huang, Xuanqiang Angelo, et al.
Veröffentlicht: (2026)
Voices of Her: Analyzing Gender Differences in the AI Publication World
von: Ding, Yiwen, et al.
Veröffentlicht: (2023)
von: Ding, Yiwen, et al.
Veröffentlicht: (2023)
CausalCite: A Causal Formulation of Paper Citations
von: Kumar, Ishan, et al.
Veröffentlicht: (2023)
von: Kumar, Ishan, et al.
Veröffentlicht: (2023)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
von: Backmann, Steffen, et al.
Veröffentlicht: (2025)
von: Backmann, Steffen, et al.
Veröffentlicht: (2025)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
ClawArena: Benchmarking AI Agents in Evolving Information Environments
von: Ji, Haonian, et al.
Veröffentlicht: (2026)
von: Ji, Haonian, et al.
Veröffentlicht: (2026)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
von: Ceraolo, Roberto, et al.
Veröffentlicht: (2024)
von: Ceraolo, Roberto, et al.
Veröffentlicht: (2024)
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
von: Piedrahita, David Guzman, et al.
Veröffentlicht: (2025)
von: Piedrahita, David Guzman, et al.
Veröffentlicht: (2025)
Natural Building Blocks for Structured World Models: Theory, Evidence, and Scaling
von: Da Costa, Lancelot, et al.
Veröffentlicht: (2025)
von: Da Costa, Lancelot, et al.
Veröffentlicht: (2025)
Causality for Natural Language Processing
von: Jin, Zhijing
Veröffentlicht: (2025)
von: Jin, Zhijing
Veröffentlicht: (2025)
Provable Privacy with Non-Private Pre-Processing
von: Hu, Yaxi, et al.
Veröffentlicht: (2024)
von: Hu, Yaxi, et al.
Veröffentlicht: (2024)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
von: Miller, Moritz, et al.
Veröffentlicht: (2026)
von: Miller, Moritz, et al.
Veröffentlicht: (2026)
CLadder: Assessing Causal Reasoning in Language Models
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
Riemannian Networks over Full-Rank Correlation Matrices
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
von: Chen, Ziheng, et al.
Veröffentlicht: (2026)
Preference Elicitation for Offline Reinforcement Learning
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
Sample-aware Adaptive Structured Pruning for Large Language Models
von: Kong, Jun, et al.
Veröffentlicht: (2025)
von: Kong, Jun, et al.
Veröffentlicht: (2025)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
von: Jenny, David F., et al.
Veröffentlicht: (2023)
von: Jenny, David F., et al.
Veröffentlicht: (2023)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
von: Draye, Florent, et al.
Veröffentlicht: (2026)
von: Draye, Florent, et al.
Veröffentlicht: (2026)
Riemannian Batch Normalization: A Gyro Approach
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
von: Chen, Ziheng, et al.
Veröffentlicht: (2025)
Counterfactual reasoning: an analysis of in-context emergence
von: Miller, Moritz, et al.
Veröffentlicht: (2025)
von: Miller, Moritz, et al.
Veröffentlicht: (2025)
Algorithmic causal structure emerging through compression
von: Wendong, Liang, et al.
Veröffentlicht: (2025)
von: Wendong, Liang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Causality can systematically address the monsters under the bench(marks)
von: Leeb, Felix, et al.
Veröffentlicht: (2025) -
Can Theoretical Physics Research Benefit from Language Agents?
von: Lu, Sirui, et al.
Veröffentlicht: (2025) -
Improving Large Language Model Safety with Contrastive Representation Learning
von: Simko, Samuel, et al.
Veröffentlicht: (2025) -
Causal Responsibility Attribution for Human-AI Collaboration
von: Qi, Yahang, et al.
Veröffentlicht: (2024) -
Causality Is Key to Understand and Balance Multiple Goals in Trustworthy ML and Foundation Models
von: Binkyte, Ruta, et al.
Veröffentlicht: (2025)