ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Penghao, Zhou, Yuhao, Wu, Mengxuan, Qin, Ziheng, Zhu, Bangyuan, Huang, Shengbin, Zhao, Xuanlei, Zhang, Panpan, Peng, Xiaojiang, Shang, Yuzhang, Yang, Jianfei, Zhu, Zheng, Chen, Tianlong, Wang, Zhangyang, Wang, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
von: Wang, Penghao, et al.
Veröffentlicht: (2025)
von: Wang, Penghao, et al.
Veröffentlicht: (2025)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
von: Cong, Wenyan, et al.
Veröffentlicht: (2025)
Deep Learning‐Based Image Recognition for Food Science and Technology: End‐to‐End Workflows and Domain‐Specific Solutions
von: Bin Liao, et al.
Veröffentlicht: (2026)
von: Bin Liao, et al.
Veröffentlicht: (2026)
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
Good Agentic Friends Do Not Just Give Verbal Advice: They Can Update Your Weights
von: Bao, Wenrui, et al.
Veröffentlicht: (2026)
von: Bao, Wenrui, et al.
Veröffentlicht: (2026)
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
von: Sun, Haoyu, et al.
Veröffentlicht: (2025)
von: Sun, Haoyu, et al.
Veröffentlicht: (2025)
Benchmarking AI Performance on End-to-End Data Science Projects
von: Hughes, Evelyn, et al.
Veröffentlicht: (2026)
von: Hughes, Evelyn, et al.
Veröffentlicht: (2026)
New Insights into Global Warming: End-to-End Visual Analysis and Prediction of Temperature Variations
von: Zhou, Meihua, et al.
Veröffentlicht: (2024)
von: Zhou, Meihua, et al.
Veröffentlicht: (2024)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification
von: Wan, Gwok-Waa, et al.
Veröffentlicht: (2025)
von: Wan, Gwok-Waa, et al.
Veröffentlicht: (2025)
MoCha:End-to-End Video Character Replacement without Structural Guidance
von: Xu, Zhengbo, et al.
Veröffentlicht: (2026)
von: Xu, Zhengbo, et al.
Veröffentlicht: (2026)
ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
von: Jin, Tengjun, et al.
Veröffentlicht: (2025)
von: Jin, Tengjun, et al.
Veröffentlicht: (2025)
End-User Training at the Amoco Research Center.
von: Kirk, Cheryl L.
Veröffentlicht: (1986)
von: Kirk, Cheryl L.
Veröffentlicht: (1986)
LLM-AutoDiff: Auto-Differentiate Any LLM Workflow
von: Yin, Li, et al.
Veröffentlicht: (2025)
von: Yin, Li, et al.
Veröffentlicht: (2025)
MDAgent: A Multi-Agent Framework for End-to-End Molecular Dynamics Research
von: Ma, Zhenyu, et al.
Veröffentlicht: (2026)
von: Ma, Zhenyu, et al.
Veröffentlicht: (2026)
WebDS: An End-to-End Benchmark for Web-based Data Science
von: Hsu, Ethan, et al.
Veröffentlicht: (2025)
von: Hsu, Ethan, et al.
Veröffentlicht: (2025)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
PRBench: End-to-end Paper Reproduction in Physics Research
von: Qiu, Shi, et al.
Veröffentlicht: (2026)
von: Qiu, Shi, et al.
Veröffentlicht: (2026)
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
GE2E-KWS: Generalized End-to-End Training and Evaluation for Zero-shot Keyword Spotting
von: Zhu, Pai, et al.
Veröffentlicht: (2024)
von: Zhu, Pai, et al.
Veröffentlicht: (2024)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
von: Tan, Qitao, et al.
Veröffentlicht: (2025)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
von: Zhu, Jiaying, et al.
Veröffentlicht: (2025)
von: Zhu, Jiaying, et al.
Veröffentlicht: (2025)
NORA: A Harness-Engineered Autonomous Research Agent for End-to-End Spatial Data Science
von: Zhou, Bing, et al.
Veröffentlicht: (2026)
von: Zhou, Bing, et al.
Veröffentlicht: (2026)
ORMind: A Cognitive-Inspired End-to-End Reasoning Framework for Operations Research
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2025)
Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge Deployment
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
von: Ji, Yuhao, et al.
Veröffentlicht: (2024)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
von: Zhu, Hongda, et al.
Veröffentlicht: (2025)
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
von: Wang, Jialing, et al.
Veröffentlicht: (2026)
von: Wang, Jialing, et al.
Veröffentlicht: (2026)
EGA-V1: Unifying Online Advertising with End-to-End Learning
von: Qiu, Junyan, et al.
Veröffentlicht: (2025)
von: Qiu, Junyan, et al.
Veröffentlicht: (2025)
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
von: Zhao, Qinyu, et al.
Veröffentlicht: (2025)
von: Zhao, Qinyu, et al.
Veröffentlicht: (2025)
Towards End-to-End Earthquake Monitoring Using a Multitask Deep Learning Model
von: Zhu, Weiqiang, et al.
Veröffentlicht: (2025)
von: Zhu, Weiqiang, et al.
Veröffentlicht: (2025)
Countering Mainstream Bias via End-to-End Adaptive Local Learning
von: Pan, Jinhao, et al.
Veröffentlicht: (2024)
von: Pan, Jinhao, et al.
Veröffentlicht: (2024)
Cool-3D: An End-to-End Thermal-Aware Framework for Early-Phase Design Space Exploration of Microfluidic-Cooled 3DICs
von: Wang, Runxi, et al.
Veröffentlicht: (2025)
von: Wang, Runxi, et al.
Veröffentlicht: (2025)
Hallucination to Consensus: Multi-Agent LLMs for End-to-End JUnit Test Generation
von: Xu, Qinghua, et al.
Veröffentlicht: (2025)
von: Xu, Qinghua, et al.
Veröffentlicht: (2025)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
Driving with A Thousand Faces: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
von: Dong, Xiaoru, et al.
Veröffentlicht: (2026)
von: Dong, Xiaoru, et al.
Veröffentlicht: (2026)
Dynamic Vision Mamba
von: Wu, Mengxuan, et al.
Veröffentlicht: (2025)
von: Wu, Mengxuan, et al.
Veröffentlicht: (2025)
From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
von: Ni, Jiliang, et al.
Veröffentlicht: (2025)
von: Ni, Jiliang, et al.
Veröffentlicht: (2025)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
von: Zhou, Sifan, et al.
Veröffentlicht: (2025)
von: Zhou, Sifan, et al.
Veröffentlicht: (2025)
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
von: Lau, Elaine, et al.
Veröffentlicht: (2026)
von: Lau, Elaine, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
von: Wang, Penghao, et al.
Veröffentlicht: (2025) -
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
von: Wang, Yuchen, et al.
Veröffentlicht: (2026) -
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
von: Cong, Wenyan, et al.
Veröffentlicht: (2025) -
Deep Learning‐Based Image Recognition for Food Science and Technology: End‐to‐End Workflows and Domain‐Specific Solutions
von: Bin Liao, et al.
Veröffentlicht: (2026) -
Benchmarking Post-Training Quantization in LLMs: Comprehensive Taxonomy, Unified Evaluation, and Comparative Analysis
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)