DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shu, Fan, Wang, Yite, Wu, Ruofan, Liu, Boyi, Yao, Zhewei, He, Yuxiong, Yan, Feng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning to Self-Evolve
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2026)
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2026)
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
Learning to Hint for Reinforcement Learning
von: Xia, Yu, et al.
Veröffentlicht: (2026)
von: Xia, Yu, et al.
Veröffentlicht: (2026)
R$^3$-SQL: Ranking Reward and Resampling for Text-to-SQL
von: Han, Hojae, et al.
Veröffentlicht: (2026)
von: Han, Hojae, et al.
Veröffentlicht: (2026)
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
von: He, Yexiao, et al.
Veröffentlicht: (2025)
von: He, Yexiao, et al.
Veröffentlicht: (2025)
ComposeRAG: A Modular and Composable RAG for Corpus-Grounded Multi-Hop Question Answering
von: Wu, Ruofan, et al.
Veröffentlicht: (2025)
von: Wu, Ruofan, et al.
Veröffentlicht: (2025)
DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models
von: Deng, Wenlong, et al.
Veröffentlicht: (2024)
von: Deng, Wenlong, et al.
Veröffentlicht: (2024)
Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL
von: Yao, Zhewei, et al.
Veröffentlicht: (2025)
von: Yao, Zhewei, et al.
Veröffentlicht: (2025)
ExCoT: Optimizing Reasoning for Text-to-SQL with Execution Feedback
von: Zhai, Bohan, et al.
Veröffentlicht: (2025)
von: Zhai, Bohan, et al.
Veröffentlicht: (2025)
DeepSpeed Data Efficiency: Improving Deep Learning Model Quality and Training Efficiency via Efficient Data Sampling and Routing
von: Li, Conglong, et al.
Veröffentlicht: (2022)
von: Li, Conglong, et al.
Veröffentlicht: (2022)
DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval
von: Sun, Maojun, et al.
Veröffentlicht: (2026)
von: Sun, Maojun, et al.
Veröffentlicht: (2026)
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2023)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
von: Yao, Shunyu, et al.
Veröffentlicht: (2024)
InFoBench: Evaluating Instruction Following Ability in Large Language Models
von: Qin, Yiwei, et al.
Veröffentlicht: (2024)
von: Qin, Yiwei, et al.
Veröffentlicht: (2024)
Agentic Verification for Ambiguous Query Disambiguation
von: Lee, Youngwon, et al.
Veröffentlicht: (2025)
von: Lee, Youngwon, et al.
Veröffentlicht: (2025)
FML-bench: Benchmarking Machine Learning Agents for Scientific Research
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
von: Zou, Qiran, et al.
Veröffentlicht: (2025)
Evaluation of Finetuned LLMs in AMR Parsing
von: Ho, Shu Han
Veröffentlicht: (2025)
von: Ho, Shu Han
Veröffentlicht: (2025)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
von: Jimenez, Carlos E., et al.
Veröffentlicht: (2023)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
von: He, Linda, et al.
Veröffentlicht: (2025)
von: He, Linda, et al.
Veröffentlicht: (2025)
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
von: Doostmohammadi, Ehsan, et al.
Veröffentlicht: (2024)
von: Doostmohammadi, Ehsan, et al.
Veröffentlicht: (2024)
Federated Data-Efficient Instruction Tuning for Large Language Models
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
von: Qin, Zhen, et al.
Veröffentlicht: (2024)
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
von: Yan, Jianhao, et al.
Veröffentlicht: (2024)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems
von: Sun, Maojun, et al.
Veröffentlicht: (2026)
von: Sun, Maojun, et al.
Veröffentlicht: (2026)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
von: Cui, Guanyu, et al.
Veröffentlicht: (2026)
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
von: Huang, Yiming, et al.
Veröffentlicht: (2024)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
von: Shu, Lei, et al.
Veröffentlicht: (2023)
von: Shu, Lei, et al.
Veröffentlicht: (2023)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
The Instruction Gap: LLMs get lost in Following Instruction
von: Tripathi, Vishesh, et al.
Veröffentlicht: (2025)
von: Tripathi, Vishesh, et al.
Veröffentlicht: (2025)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
von: Sun, Wangtao, et al.
Veröffentlicht: (2024)
Thinking LLMs: General Instruction Following with Thought Generation
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2024)
von: Bajpai, Prasoon, et al.
Veröffentlicht: (2024)
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
von: Han, Hojae, et al.
Veröffentlicht: (2025)
von: Han, Hojae, et al.
Veröffentlicht: (2025)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning to Self-Evolve
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2026) -
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026) -
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024) -
Learning to Hint for Reinforcement Learning
von: Xia, Yu, et al.
Veröffentlicht: (2026) -
R$^3$-SQL: Ranking Reward and Resampling for Text-to-SQL
von: Han, Hojae, et al.
Veröffentlicht: (2026)