SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Ken, Bhat, Advait, Merrill, Mike A, West, Robert, Liu, Xin, McDuff, Daniel, Althoff, Tim |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Are the Odds? Language Models Are Capable of Probabilistic Reasoning
by: Paruchuri, Akshay, et al.
Published: (2024)
by: Paruchuri, Akshay, et al.
Published: (2024)
Language Models Still Struggle to Zero-shot Reason about Time Series
by: Merrill, Mike A., et al.
Published: (2024)
by: Merrill, Mike A., et al.
Published: (2024)
Substance over Style: Evaluating Proactive Conversational Coaching Agents
by: Srinivas, Vidya, et al.
Published: (2025)
by: Srinivas, Vidya, et al.
Published: (2025)
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025)
by: Kim, Yubin, et al.
Published: (2025)
Transforming Wearable Data into Personal Health Insights using Large Language Model Agents
by: Merrill, Mike A., et al.
Published: (2024)
by: Merrill, Mike A., et al.
Published: (2024)
Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
RADAR: Benchmarking Language Models on Imperfect Tabular Data
by: Gu, Ken, et al.
Published: (2025)
by: Gu, Ken, et al.
Published: (2025)
The chemistry of interstitial waters at DSDP Site 45-395
by: McDuff, Russell E
Published: (1984)
by: McDuff, Russell E
Published: (1984)
(Table 3) Chemistry in water at DSDP Site 45-395
by: McDuff, Russell E
Published: (1984)
by: McDuff, Russell E
Published: (1984)
(Table 2) Silicon and nitrate concentrations in water samples at DSDP Hole 45-395A
by: McDuff, Russell E
Published: (1984)
by: McDuff, Russell E
Published: (1984)
(Table 3) Interstitial water elemental composition at DSDP Leg 86 Holes
by: McDuff, Russell E
Published: (1985)
by: McDuff, Russell E
Published: (1985)
BLADE: Benchmarking Language Model Agents for Data-Driven Science
by: Gu, Ken, et al.
Published: (2024)
by: Gu, Ken, et al.
Published: (2024)
Disentangling Reasoning and Knowledge in Medical Large Language Models
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning
by: Merrill, Scott, et al.
Published: (2026)
by: Merrill, Scott, et al.
Published: (2026)
KoLA: Carefully Benchmarking World Knowledge of Large Language Models
by: Yu, Jifan, et al.
Published: (2023)
by: Yu, Jifan, et al.
Published: (2023)
Are Language Models Actually Useful for Time Series Forecasting?
by: Tan, Mingtian, et al.
Published: (2024)
by: Tan, Mingtian, et al.
Published: (2024)
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
by: McDuff, Daniel, et al.
Published: (2025)
by: McDuff, Daniel, et al.
Published: (2025)
Polyfold fundamental classes and globally structured multivalued perturbations
by: McDuff, Dusa, et al.
Published: (2024)
by: McDuff, Dusa, et al.
Published: (2024)
Sesquicuspidal curves, scattering diagrams, and symplectic nonsqueezing
by: McDuff, Dusa, et al.
Published: (2024)
by: McDuff, Dusa, et al.
Published: (2024)
Singular algebraic curves and infinite symplectic staircases
by: McDuff, Dusa, et al.
Published: (2024)
by: McDuff, Dusa, et al.
Published: (2024)
Symplectic capacities, unperturbed curves, and convex toric domains
by: McDuff, Dusa, et al.
Published: (2021)
by: McDuff, Dusa, et al.
Published: (2021)
ReaSeq: Unleashing World Knowledge via Reasoning for Sequential Modeling
by: Tang, Jiakai, et al.
Published: (2025)
by: Tang, Jiakai, et al.
Published: (2025)
A Demonstration of Adaptive Collaboration of Large Language Models for Medical Decision-Making
by: Kim, Yubin, et al.
Published: (2024)
by: Kim, Yubin, et al.
Published: (2024)
SensorLM: Learning the Language of Wearable Sensors
by: Zhang, Yuwei, et al.
Published: (2025)
by: Zhang, Yuwei, et al.
Published: (2025)
LogiNumSynth: Synthesizing Joint Logical-Numerical Reasoning Problems for Language Models
by: Liu, Yiwei, et al.
Published: (2025)
by: Liu, Yiwei, et al.
Published: (2025)
Towards Localized and Disentangled Knowledge Editing for Multimodal Large Language Models
by: Gu, Leijiang, et al.
Published: (2026)
by: Gu, Leijiang, et al.
Published: (2026)
Language and Knowledge of the World
by: Paivio, Allan
Published: (1974)
by: Paivio, Allan
Published: (1974)
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
by: Arčon, Tjaša, et al.
Published: (2026)
by: Arčon, Tjaša, et al.
Published: (2026)
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
by: Jalilian, Laleh, et al.
Published: (2024)
by: Jalilian, Laleh, et al.
Published: (2024)
How Suboptimal is Training rPPG Models with Videos and Targets from Different Body Sites?
by: Braun, Björn, et al.
Published: (2024)
by: Braun, Björn, et al.
Published: (2024)
Passive Measurement of Autonomic Arousal in Real-World Settings
by: Abdel-Ghaffar, Samy, et al.
Published: (2025)
by: Abdel-Ghaffar, Samy, et al.
Published: (2025)
Improving Neural Question Generation using World Knowledge
by: Gupta, Deepak, et al.
Published: (2019)
by: Gupta, Deepak, et al.
Published: (2019)
Carpe Diem: On the Evaluation of World Knowledge in Lifelong Language Models
by: Kim, Yujin, et al.
Published: (2023)
by: Kim, Yujin, et al.
Published: (2023)
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
by: Veselovsky, Veniamin, et al.
Published: (2025)
by: Veselovsky, Veniamin, et al.
Published: (2025)
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
by: Moskovskiy, Daniil, et al.
Published: (2025)
by: Moskovskiy, Daniil, et al.
Published: (2025)
Shiksha: A Technical Domain focused Translation Dataset and Model for Indian Languages
by: Joglekar, Advait, et al.
Published: (2024)
by: Joglekar, Advait, et al.
Published: (2024)
Leveraging Test Driven Development with Large Language Models for Reliable and Verifiable Spreadsheet Code Generation: A Research Framework
by: Thorne, Simon, et al.
Published: (2025)
by: Thorne, Simon, et al.
Published: (2025)
Unmask It! AI-Generated Product Review Detection in Dravidian Languages
by: De, Somsubhra, et al.
Published: (2025)
by: De, Somsubhra, et al.
Published: (2025)
Making Large Language Models into World Models with Precondition and Effect Knowledge
by: Xie, Kaige, et al.
Published: (2024)
by: Xie, Kaige, et al.
Published: (2024)
Follow the Path: Reasoning over Knowledge Graph Paths to Improve Large Language Model Factuality
by: Zhang, Mike, et al.
Published: (2025)
by: Zhang, Mike, et al.
Published: (2025)
Similar Items
-
What Are the Odds? Language Models Are Capable of Probabilistic Reasoning
by: Paruchuri, Akshay, et al.
Published: (2024) -
Language Models Still Struggle to Zero-shot Reason about Time Series
by: Merrill, Mike A., et al.
Published: (2024) -
Substance over Style: Evaluating Proactive Conversational Coaching Agents
by: Srinivas, Vidya, et al.
Published: (2025) -
InvThink: Premortem Reasoning for Safer Language Models
by: Kim, Yubin, et al.
Published: (2025) -
Transforming Wearable Data into Personal Health Insights using Large Language Model Agents
by: Merrill, Mike A., et al.
Published: (2024)