Gespeichert in:
| Hauptverfasser: | Zhao, Xinran, Zheng, Boyuan, Si, Chenglei, Yu, Haofei, Liu, Ken, Zhou, Runlong, Li, Ruochen, Chen, Tong, Li, Xiang, Zhang, Yiming, Wu, Tongshuang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2508.19200 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
von: Si, Chenglei, et al.
Veröffentlicht: (2025)
Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
von: Zhao, Xinran, et al.
Veröffentlicht: (2024)
von: Zhao, Xinran, et al.
Veröffentlicht: (2024)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
von: Si, Chenglei, et al.
Veröffentlicht: (2023)
von: Si, Chenglei, et al.
Veröffentlicht: (2023)
Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models
von: Zhao, Xinran, et al.
Veröffentlicht: (2024)
von: Zhao, Xinran, et al.
Veröffentlicht: (2024)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
von: Zeng, Yixiao, et al.
Veröffentlicht: (2025)
von: Zeng, Yixiao, et al.
Veröffentlicht: (2025)
Revela: Dense Retriever Learning via Language Modeling
von: Cai, Fengyu, et al.
Veröffentlicht: (2025)
von: Cai, Fengyu, et al.
Veröffentlicht: (2025)
MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
von: Kalra, Jushaan Singh, et al.
Veröffentlicht: (2025)
von: Kalra, Jushaan Singh, et al.
Veröffentlicht: (2025)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
von: Zhang, Yilin, et al.
Veröffentlicht: (2025)
von: Zhang, Yilin, et al.
Veröffentlicht: (2025)
SPHERE: An Evaluation Card for Human-AI Systems
von: Ma, Qianou, et al.
Veröffentlicht: (2025)
von: Ma, Qianou, et al.
Veröffentlicht: (2025)
Towards Execution-Grounded Automated AI Research
von: Si, Chenglei, et al.
Veröffentlicht: (2026)
von: Si, Chenglei, et al.
Veröffentlicht: (2026)
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
von: Zhao, Ruochen, et al.
Veröffentlicht: (2024)
von: Zhao, Ruochen, et al.
Veröffentlicht: (2024)
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
von: Zhou, Runlong, et al.
Veröffentlicht: (2024)
von: Zhou, Runlong, et al.
Veröffentlicht: (2024)
Completion $\neq$ Collaboration: Scaling Collaborative Effort with Agents
von: Shen, Shannon Zejiang, et al.
Veröffentlicht: (2025)
von: Shen, Shannon Zejiang, et al.
Veröffentlicht: (2025)
Improving Attributed Long-form Question Answering with Intent Awareness
von: Zhao, Xinran, et al.
Veröffentlicht: (2026)
von: Zhao, Xinran, et al.
Veröffentlicht: (2026)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation
von: Lei, Xinping, et al.
Veröffentlicht: (2025)
von: Lei, Xinping, et al.
Veröffentlicht: (2025)
CASCADE Your Datasets for Cross-Mode Knowledge Retrieval of Language Models
von: Zhou, Runlong, et al.
Veröffentlicht: (2025)
von: Zhou, Runlong, et al.
Veröffentlicht: (2025)
Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation
von: Li, Tong, et al.
Veröffentlicht: (2025)
von: Li, Tong, et al.
Veröffentlicht: (2025)
Think out Loud: Emotion Deducing Explanation in Dialogues
von: Li, Jiangnan, et al.
Veröffentlicht: (2024)
von: Li, Jiangnan, et al.
Veröffentlicht: (2024)
DR-Arena: an Automated Evaluation Framework for Deep Research Agents
von: Gao, Yiwen, et al.
Veröffentlicht: (2026)
von: Gao, Yiwen, et al.
Veröffentlicht: (2026)
ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
von: Li, Sunzhu, et al.
Veröffentlicht: (2025)
von: Li, Sunzhu, et al.
Veröffentlicht: (2025)
Evaluating Mathematical Reasoning Beyond Accuracy
von: Xia, Shijie, et al.
Veröffentlicht: (2024)
von: Xia, Shijie, et al.
Veröffentlicht: (2024)
Contextual Experience Replay for Self-Improvement of Language Agents
von: Liu, Yitao, et al.
Veröffentlicht: (2025)
von: Liu, Yitao, et al.
Veröffentlicht: (2025)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
von: Zhao, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chenyang, et al.
Veröffentlicht: (2024)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
von: Deng, Wenhao, et al.
Veröffentlicht: (2025)
The Crucial Role of Samplers in Online Direct Preference Optimization
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
von: Shi, Ruizhe, et al.
Veröffentlicht: (2024)
AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
von: Wu, Jinyang, et al.
Veröffentlicht: (2025)
Multi-Agent Causal Reasoning for Suicide Ideation Detection Through Online Conversations
von: Li, Jun, et al.
Veröffentlicht: (2026)
von: Li, Jun, et al.
Veröffentlicht: (2026)
Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
von: Yong, Xixian, et al.
Veröffentlicht: (2025)
von: Yong, Xixian, et al.
Veröffentlicht: (2025)
Checklists Are Better Than Reward Models For Aligning Language Models
von: Viswanathan, Vijay, et al.
Veröffentlicht: (2025)
von: Viswanathan, Vijay, et al.
Veröffentlicht: (2025)
From Handcrafted Features to LLMs: A Brief Survey for Machine Translation Quality Estimation
von: Zhao, Haofei, et al.
Veröffentlicht: (2024)
von: Zhao, Haofei, et al.
Veröffentlicht: (2024)
LiveTradeBench: Seeking Real-World Alpha with Large Language Models
von: Yu, Haofei, et al.
Veröffentlicht: (2025)
von: Yu, Haofei, et al.
Veröffentlicht: (2025)
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
von: Liu, Yiren, et al.
Veröffentlicht: (2025)
von: Liu, Yiren, et al.
Veröffentlicht: (2025)
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
von: Jian, Yichang, et al.
Veröffentlicht: (2026)
von: Jian, Yichang, et al.
Veröffentlicht: (2026)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
MLR-Copilot: Autonomous Machine Learning Research based on Large Language Models Agents
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
von: Li, Ruochen, et al.
Veröffentlicht: (2024)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
von: Gandhi, Saumya, et al.
Veröffentlicht: (2024)
von: Gandhi, Saumya, et al.
Veröffentlicht: (2024)
"I didn't Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration
von: Kim, Eunsu, et al.
Veröffentlicht: (2026)
von: Kim, Eunsu, et al.
Veröffentlicht: (2026)
Effectively Controlling Reasoning Models through Thinking Intervention
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
von: Si, Chenglei, et al.
Veröffentlicht: (2025) -
Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
von: Zhao, Xinran, et al.
Veröffentlicht: (2024) -
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
von: Si, Chenglei, et al.
Veröffentlicht: (2023) -
Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models
von: Zhao, Xinran, et al.
Veröffentlicht: (2024) -
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
von: Zeng, Yixiao, et al.
Veröffentlicht: (2025)