NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Zheng, Tianshi, Tam, Kelvin Kiu-Wai, Nguyen, Newt Hue-Nam K., Xu, Baixuan, Wang, Zhaowei, Cheng, Jiayang, Tsang, Hong Ting, Wang, Weiqi, Bai, Jiaxin, Fang, Tianqing, Song, Yangqiu, Wong, Ginny Y., See, Simon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning
por: Zheng, Tianshi, et al.
Publicado: (2026)
por: Zheng, Tianshi, et al.
Publicado: (2026)
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
por: Xu, Baixuan, et al.
Publicado: (2025)
por: Xu, Baixuan, et al.
Publicado: (2025)
From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
por: Zheng, Tianshi, et al.
Publicado: (2025)
por: Zheng, Tianshi, et al.
Publicado: (2025)
CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense Reasoning
por: Wang, Weiqi, et al.
Publicado: (2024)
por: Wang, Weiqi, et al.
Publicado: (2024)
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
por: Zheng, Tianshi, et al.
Publicado: (2025)
por: Zheng, Tianshi, et al.
Publicado: (2025)
On the Role of Entity and Event Level Conceptualization in Generalizable Reasoning: A Survey of Tasks, Methods, Applications, and Future Directions
por: Wang, Weiqi, et al.
Publicado: (2024)
por: Wang, Weiqi, et al.
Publicado: (2024)
Legal Rule Induction: Towards Generalizable Principle Discovery from Analogous Judicial Precedents
por: Fan, Wei, et al.
Publicado: (2025)
por: Fan, Wei, et al.
Publicado: (2025)
KnowShiftQA: How Robust are RAG Systems when Textbook Knowledge Shifts in K-12 Education?
por: Zheng, Tianshi, et al.
Publicado: (2024)
por: Zheng, Tianshi, et al.
Publicado: (2024)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
por: Zong, Qing, et al.
Publicado: (2025)
por: Zong, Qing, et al.
Publicado: (2025)
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
por: Zheng, Tianshi, et al.
Publicado: (2025)
por: Zheng, Tianshi, et al.
Publicado: (2025)
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
por: Ren, Xiyu, et al.
Publicado: (2026)
por: Ren, Xiyu, et al.
Publicado: (2026)
AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
por: Tsang, Hong Ting, et al.
Publicado: (2025)
por: Tsang, Hong Ting, et al.
Publicado: (2025)
AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility Estimation
por: Wang, Zhaowei, et al.
Publicado: (2024)
por: Wang, Zhaowei, et al.
Publicado: (2024)
Patterns Over Principles: The Fragility of Inductive Reasoning in LLMs under Noisy Observations
por: Li, Chunyang, et al.
Publicado: (2025)
por: Li, Chunyang, et al.
Publicado: (2025)
Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization
por: He, Mutian, et al.
Publicado: (2022)
por: He, Mutian, et al.
Publicado: (2022)
ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning
por: Zhang, Liyu, et al.
Publicado: (2024)
por: Zhang, Liyu, et al.
Publicado: (2024)
EcomEdit: An Automated E-commerce Knowledge Editing Framework for Enhanced Product and Purchase Intention Understanding
por: Lau, Ching Ming Samuel, et al.
Publicado: (2024)
por: Lau, Ching Ming Samuel, et al.
Publicado: (2024)
CKBP v2: Better Annotation and Reasoning for Commonsense Knowledge Base Population
por: Fang, Tianqing, et al.
Publicado: (2023)
por: Fang, Tianqing, et al.
Publicado: (2023)
ChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations
por: Chan, Chunkit, et al.
Publicado: (2023)
por: Chan, Chunkit, et al.
Publicado: (2023)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
por: Wang, Zhaowei, et al.
Publicado: (2025)
por: Wang, Zhaowei, et al.
Publicado: (2025)
CLR-Fact: Evaluating the Complex Logical Reasoning Capability of Large Language Models over Factual Knowledge
por: Zheng, Tianshi, et al.
Publicado: (2024)
por: Zheng, Tianshi, et al.
Publicado: (2024)
ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty
por: Zong, Qing, et al.
Publicado: (2024)
por: Zong, Qing, et al.
Publicado: (2024)
Enhancing Transformers for Generalizable First-Order Logical Entailment
por: Zheng, Tianshi, et al.
Publicado: (2025)
por: Zheng, Tianshi, et al.
Publicado: (2025)
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
por: Wang, Zhaowei, et al.
Publicado: (2023)
por: Wang, Zhaowei, et al.
Publicado: (2023)
Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
por: Xu, Baixuan, et al.
Publicado: (2025)
por: Xu, Baixuan, et al.
Publicado: (2025)
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling
por: Shi, Haochen, et al.
Publicado: (2025)
por: Shi, Haochen, et al.
Publicado: (2025)
Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction
por: Han, Kaiqiao, et al.
Publicado: (2024)
por: Han, Kaiqiao, et al.
Publicado: (2024)
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
por: Do, Quyet V., et al.
Publicado: (2024)
por: Do, Quyet V., et al.
Publicado: (2024)
Towards Subgraph Isomorphism Counting with Graph Kernels
por: Liu, Xin, et al.
Publicado: (2024)
por: Liu, Xin, et al.
Publicado: (2024)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
por: Wang, Weiqi, et al.
Publicado: (2024)
por: Wang, Weiqi, et al.
Publicado: (2024)
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
por: Fang, Tianqing, et al.
Publicado: (2023)
por: Fang, Tianqing, et al.
Publicado: (2023)
$\mathbb{R}^{2k}$ is Theoretically Large Enough for Embedding-based Top-$k$ Retrieval
por: Wang, Zihao, et al.
Publicado: (2026)
por: Wang, Zihao, et al.
Publicado: (2026)
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
por: Mo, Yunxiang, et al.
Publicado: (2025)
por: Mo, Yunxiang, et al.
Publicado: (2025)
Structuring the Unstructured: A Systematic Review of Text-to-Structure Generation for Agentic AI with a Universal Evaluation Framework
por: Deng, Zheye, et al.
Publicado: (2025)
por: Deng, Zheye, et al.
Publicado: (2025)
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
por: Fan, Wei, et al.
Publicado: (2024)
por: Fan, Wei, et al.
Publicado: (2024)
Persona Knowledge-Aligned Prompt Tuning Method for Online Debate
por: Chan, Chunkit, et al.
Publicado: (2024)
por: Chan, Chunkit, et al.
Publicado: (2024)
Advancing Abductive Reasoning in Knowledge Graphs through Complex Logical Hypothesis Generation
por: Bai, Jiaxin, et al.
Publicado: (2023)
por: Bai, Jiaxin, et al.
Publicado: (2023)
EntailE: Introducing Textual Entailment in Commonsense Knowledge Graph Completion
por: Su, Ying, et al.
Publicado: (2024)
por: Su, Ying, et al.
Publicado: (2024)
MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding
por: Xu, Baixuan, et al.
Publicado: (2024)
por: Xu, Baixuan, et al.
Publicado: (2024)
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
por: Zheng, Zihao, et al.
Publicado: (2024)
por: Zheng, Zihao, et al.
Publicado: (2024)
Ejemplares similares
-
SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning
por: Zheng, Tianshi, et al.
Publicado: (2026) -
The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
por: Xu, Baixuan, et al.
Publicado: (2025) -
From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
por: Zheng, Tianshi, et al.
Publicado: (2025) -
CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense Reasoning
por: Wang, Weiqi, et al.
Publicado: (2024) -
LogiDynamics: Unraveling the Dynamics of Inductive, Abductive and Deductive Logical Inferences in LLM Reasoning
por: Zheng, Tianshi, et al.
Publicado: (2025)