RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xinnuo, Lawrence, Rachel, Dubey, Kshitij, Pandey, Atharva, Ueno, Risa, Falck, Fabian, Nori, Aditya V., Sharma, Rahul, Sharma, Amit, Gonzalez, Javier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
by: Pandey, Atharva, et al.
Published: (2025)
by: Pandey, Atharva, et al.
Published: (2025)
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Compositional Causal Reasoning Evaluation in Language Models
by: Maasch, Jacqueline R. M. A., et al.
Published: (2025)
by: Maasch, Jacqueline R. M. A., et al.
Published: (2025)
Reasoning Elicitation in Language Models via Counterfactual Feedback
by: Hüyük, Alihan, et al.
Published: (2024)
by: Hüyük, Alihan, et al.
Published: (2024)
Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language Models
by: González, Javier, et al.
Published: (2024)
by: González, Javier, et al.
Published: (2024)
Equivalence Checking of ML GPU Kernels
by: Dubey, Kshitij, et al.
Published: (2025)
by: Dubey, Kshitij, et al.
Published: (2025)
Teaching Transformers Causal Reasoning through Axiomatic Training
by: Vashishtha, Aniket, et al.
Published: (2024)
by: Vashishtha, Aniket, et al.
Published: (2024)
A Critical Review of Causal Reasoning Benchmarks for Large Language Models
by: Yang, Linying, et al.
Published: (2024)
by: Yang, Linying, et al.
Published: (2024)
ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis
by: Naik, Atharva, et al.
Published: (2026)
by: Naik, Atharva, et al.
Published: (2026)
RTMS: A Real-Time Multimodal Scaffolding System for Improving Debugging in Computing Education
by: Golrang, Anahita, et al.
Published: (2026)
by: Golrang, Anahita, et al.
Published: (2026)
Can providing feedback on gaze and mental-effort synchrony improve pair programming performance?
by: Golrang, Anahita, et al.
Published: (2026)
by: Golrang, Anahita, et al.
Published: (2026)
Cognitive Alignment Drives Attention: Modeling and Supporting Socially Shared Regulation in Pair Programming
by: Golrang, Anahita, et al.
Published: (2026)
by: Golrang, Anahita, et al.
Published: (2026)
IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
by: Zhang, Xikai, et al.
Published: (2025)
by: Zhang, Xikai, et al.
Published: (2025)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)
by: Maheshwari, Aditya, et al.
Published: (2026)
ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
by: Banyas, Peter, et al.
Published: (2025)
by: Banyas, Peter, et al.
Published: (2025)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
by: Lei, Jiayi, et al.
Published: (2025)
by: Lei, Jiayi, et al.
Published: (2025)
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
Automated Defect Identification and Categorization in NDE 4.0 with the Application of Artificial Intelligence
by: Sharma, Aditya
Published: (2025)
by: Sharma, Aditya
Published: (2025)
Experiential Marketing Research Dataset (2005–2024): Scopus-Based Bibliometric Data
by: Aditya, Sharma
Published: (2025)
by: Aditya, Sharma
Published: (2025)
Studies on Carrollian Quantum Field Theories
by: Sharma, Aditya
Published: (2025)
by: Sharma, Aditya
Published: (2025)
Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks
by: Sharma, Arun
Published: (2026)
by: Sharma, Arun
Published: (2026)
Synthesis of Copper Oxide (CuO) Nanofillers for Green Energy Applications
by: Kailash Kumar, et al.
Published: (2026)
by: Kailash Kumar, et al.
Published: (2026)
Mental health of computing professionals and students: A systematic literature review
by: Takaoka, Alicia Julia Wilson, et al.
Published: (2024)
by: Takaoka, Alicia Julia Wilson, et al.
Published: (2024)
ProPACT: A Proactive AI-Driven Adaptive Collaborative Tutor for Pair Programming
by: Golrang, Anahita, et al.
Published: (2026)
by: Golrang, Anahita, et al.
Published: (2026)
Quantum chaos in PT symmetric quantum systems
by: Sharma, Kshitij, et al.
Published: (2024)
by: Sharma, Kshitij, et al.
Published: (2024)
Fatha-2
by: Sharma, Amit
Published: (2025)
by: Sharma, Amit
Published: (2025)
AI Accelerators for Large Language Model Inference: Architecture Analysis and Scaling Strategies
by: Sharma, Amit
Published: (2025)
by: Sharma, Amit
Published: (2025)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
by: Vashistha, Sachin, et al.
Published: (2025)
by: Vashistha, Sachin, et al.
Published: (2025)
A Fourier Space Perspective on Diffusion Models
by: Falck, Fabian, et al.
Published: (2025)
by: Falck, Fabian, et al.
Published: (2025)
Intramolecular C─H (Hetero)arylation of Histidine and Its Derivatives: Synthesis of a New Structural Class of Heteroaromatic Amino Acids
by: Kamya Rao, et al.
Published: (2026)
by: Kamya Rao, et al.
Published: (2026)
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
by: Mundada, Gagan, et al.
Published: (2025)
by: Mundada, Gagan, et al.
Published: (2025)
Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing
by: Mutsaddi, Atharva, et al.
Published: (2025)
by: Mutsaddi, Atharva, et al.
Published: (2025)
Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts
by: Ochieng, Millicent, et al.
Published: (2025)
by: Ochieng, Millicent, et al.
Published: (2025)
IMAGINE: Intelligent Multi-Agent Godot-based Indoor Networked Exploration
by: Leite, Tiago, et al.
Published: (2026)
by: Leite, Tiago, et al.
Published: (2026)
Governance in the age of artificial intelligence: A comparative analysis of policy framework in BRICS nations
by: Animesh Kumar Sharma, et al.
Published: (2025)
by: Animesh Kumar Sharma, et al.
Published: (2025)
Cautionary Tales on Synthetic Controls in Survival Analyses
by: Curth, Alicia, et al.
Published: (2023)
by: Curth, Alicia, et al.
Published: (2023)
A Physical Analogy between Molecular Ordering and SAT-to-Ising Annealing
by: Dubey, ShivKishan, et al.
Published: (2025)
by: Dubey, ShivKishan, et al.
Published: (2025)
Optimized Gradient Tracking for Decentralized Online Learning
by: Sharma, Shivangi Dubey, et al.
Published: (2023)
by: Sharma, Shivangi Dubey, et al.
Published: (2023)
SciNets: Graph-Constrained Multi-Hop Reasoning for Scientific Literature Synthesis
by: Dubey, Sauhard
Published: (2025)
by: Dubey, Sauhard
Published: (2025)
Synthesis of 8‐Hydroxy Quinoline Derivatives: A Revolutionary Target in Medicinal Chemistry (2020–2025)
by: Radhika Sharma, et al.
Published: (2025)
by: Radhika Sharma, et al.
Published: (2025)
Similar Items
-
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
by: Pandey, Atharva, et al.
Published: (2025) -
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
by: Lin, Victoria, et al.
Published: (2026) -
Compositional Causal Reasoning Evaluation in Language Models
by: Maasch, Jacqueline R. M. A., et al.
Published: (2025) -
Reasoning Elicitation in Language Models via Counterfactual Feedback
by: Hüyük, Alihan, et al.
Published: (2024) -
Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language Models
by: González, Javier, et al.
Published: (2024)