AI Can Learn Scientific Taste
Fuente:
arXiv
Saved in:
| Main Authors: | Tong, Jingqi, Li, Mingzhe, Li, Hangcheng, Yang, Yongzhuo, Mou, Yurong, Ma, Weijie, Xi, Zhiheng, Chen, Hongji, Liu, Xiaoran, Cheng, Qinyuan, Zhang, Ming, Chen, Qiguang, Ge, Weifeng, Guo, Qipeng, Ying, Tianlei, Sun, Tianxiang, Zheng, Yining, Chen, Xinchi, Zhao, Jun, Ding, Ning, Huang, Xuanjing, Jiang, Yugang, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
by: Fang, Shicheng, et al.
Published: (2026)
by: Fang, Shicheng, et al.
Published: (2026)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
by: Wang, Weiyi, et al.
Published: (2026)
by: Wang, Weiyi, et al.
Published: (2026)
Agent Alignment in Evolving Social Norms
by: Li, Shimin, et al.
Published: (2024)
by: Li, Shimin, et al.
Published: (2024)
How to Mitigate Overfitting in Weak-to-strong Generalization?
by: Shi, Junhao, et al.
Published: (2025)
by: Shi, Junhao, et al.
Published: (2025)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
by: Zhao, Jun, et al.
Published: (2024)
by: Zhao, Jun, et al.
Published: (2024)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
by: Tong, Jingqi, et al.
Published: (2025)
by: Tong, Jingqi, et al.
Published: (2025)
DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
by: Zeng, Zhiyuan, et al.
Published: (2024)
by: Zeng, Zhiyuan, et al.
Published: (2024)
Scaling Laws for Fact Memorization of Large Language Models
by: Lu, Xingyu, et al.
Published: (2024)
by: Lu, Xingyu, et al.
Published: (2024)
LLM can Achieve Self-Regulation via Hyperparameter Aware Generation
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
by: Yin, Zhangyue, et al.
Published: (2024)
by: Yin, Zhangyue, et al.
Published: (2024)
Dynamic and Generalizable Process Reward Modeling
by: Yin, Zhangyue, et al.
Published: (2025)
by: Yin, Zhangyue, et al.
Published: (2025)
Model Utility Law: Evaluating LLMs beyond Performance through Mechanism Interpretable Metric
by: Cao, Yixin, et al.
Published: (2025)
by: Cao, Yixin, et al.
Published: (2025)
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Structure-enhanced Contrastive Learning for Graph Clustering
by: Wu, Xunlian, et al.
Published: (2024)
by: Wu, Xunlian, et al.
Published: (2024)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
by: Wang, Yuxin, et al.
Published: (2025)
by: Wang, Yuxin, et al.
Published: (2025)
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
by: Yang, Jie, et al.
Published: (2025)
by: Yang, Jie, et al.
Published: (2025)
Case2Code: Scalable Synthetic Data for Code Generation
by: Shao, Yunfan, et al.
Published: (2024)
by: Shao, Yunfan, et al.
Published: (2024)
Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding
by: Chen, Tianlei, et al.
Published: (2025)
by: Chen, Tianlei, et al.
Published: (2025)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
by: Wang, Yubang, et al.
Published: (2026)
by: Wang, Yubang, et al.
Published: (2026)
ARISE: An Adaptive Resolution-Aware Metric for Test-Time Scaling Evaluation in Large Reasoning Models
by: Yin, Zhangyue, et al.
Published: (2025)
by: Yin, Zhangyue, et al.
Published: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
by: Song, Yuerong, et al.
Published: (2025)
by: Song, Yuerong, et al.
Published: (2025)
Design Strategies for Laser Additive Manufacturing of High‐Performance Alloys With Uniform Mechanical Properties
by: Jingqi Zhang, et al.
Published: (2026)
by: Jingqi Zhang, et al.
Published: (2026)
LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
by: Lv, Kai, et al.
Published: (2024)
by: Lv, Kai, et al.
Published: (2024)
Improve Cross-domain Mixed Sampling with Guidance Training for Adaptive Segmentation
by: Zhou, Wenlve, et al.
Published: (2024)
by: Zhou, Wenlve, et al.
Published: (2024)
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
by: Wang, Yikun, et al.
Published: (2025)
by: Wang, Yikun, et al.
Published: (2025)
The Combination of Video and Virtual Experiments in Online Scientific Inquiry: Effects of Learning Sequence and Representational Integration Scaffold
by: Hui Chen, et al.
Published: (2026)
by: Hui Chen, et al.
Published: (2026)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Late Fusion Multi-task Learning for Semiparametric Inference with Nuisance Parameters
by: Bhattacharya, Sohom, et al.
Published: (2025)
by: Bhattacharya, Sohom, et al.
Published: (2025)
Unified Active Retrieval for Retrieval Augmented Generation
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
by: Chen, Qiaosheng, et al.
Published: (2025)
by: Chen, Qiaosheng, et al.
Published: (2025)
Similar Items
-
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
by: Tong, Jingqi, et al.
Published: (2025) -
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
by: Fang, Shicheng, et al.
Published: (2026) -
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026) -
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
by: Wang, Weiyi, et al.
Published: (2026) -
Agent Alignment in Evolving Social Norms
by: Li, Shimin, et al.
Published: (2024)