A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Qingchuan, Ma, Yuexiao, Xie, Yongkang, Xie, Tianyu, Zheng, Xiawu, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ALGOGEN: Tool-Generated Verifiable Traces for Reliable Algorithm Visualization
by: Liao, Kunpeng, et al.
Published: (2026)
by: Liao, Kunpeng, et al.
Published: (2026)
Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
by: Ma, Qingchuan, et al.
Published: (2025)
by: Ma, Qingchuan, et al.
Published: (2025)
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
by: Zheng, Xuzhe, et al.
Published: (2026)
by: Zheng, Xuzhe, et al.
Published: (2026)
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
by: Chen, Luoxin, et al.
Published: (2026)
by: Chen, Luoxin, et al.
Published: (2026)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
by: Xie, Tianyu, et al.
Published: (2026)
by: Xie, Tianyu, et al.
Published: (2026)
Towards Verified and Targeted Explanations through Formal Methods
by: Wang, Hanchen David, et al.
Published: (2026)
by: Wang, Hanchen David, et al.
Published: (2026)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
by: Yu, Zhouliang, et al.
Published: (2025)
by: Yu, Zhouliang, et al.
Published: (2025)
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models
by: Oh, Jio, et al.
Published: (2024)
by: Oh, Jio, et al.
Published: (2024)
Polybasic Speculative Decoding Through a Theoretical Perspective
by: Wang, Ruilin, et al.
Published: (2025)
by: Wang, Ruilin, et al.
Published: (2025)
Distribution Fitting for Combating Mode Collapse in Generative Adversarial Networks
by: Gong, Yanxiang, et al.
Published: (2022)
by: Gong, Yanxiang, et al.
Published: (2022)
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
by: Gull, Ayesha, et al.
Published: (2025)
by: Gull, Ayesha, et al.
Published: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
by: Guo, Song, et al.
Published: (2024)
by: Guo, Song, et al.
Published: (2024)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
Proving the Coding Interview: A Benchmark for Formally Verified Code Generation
by: Dougherty, Quinn, et al.
Published: (2025)
by: Dougherty, Quinn, et al.
Published: (2025)
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
by: Zhang, Xiyuan, et al.
Published: (2026)
by: Zhang, Xiyuan, et al.
Published: (2026)
FAME: Formal Abstract Minimal Explanation for Neural Networks
by: Boumazouza, Ryma, et al.
Published: (2026)
by: Boumazouza, Ryma, et al.
Published: (2026)
CLEVER: A Curated Benchmark for Formally Verified Code Generation
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
by: Xie, Zhuohan, et al.
Published: (2025)
by: Xie, Zhuohan, et al.
Published: (2025)
AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
by: Aggarwal, Pranjal, et al.
Published: (2024)
by: Aggarwal, Pranjal, et al.
Published: (2024)
Formally Verifying and Explaining Sepsis Treatment Policies with COOL-MC
by: Gross, Dennis
Published: (2026)
by: Gross, Dennis
Published: (2026)
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
by: Guan, Xinyan, et al.
Published: (2024)
by: Guan, Xinyan, et al.
Published: (2024)
MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
by: Yousefzadeh, Roozbeh, et al.
Published: (2025)
by: Yousefzadeh, Roozbeh, et al.
Published: (2025)
Talking with Verifiers: Automatic Specification Generation for Neural Network Verification
by: Elboher, Yizhak Y., et al.
Published: (2026)
by: Elboher, Yizhak Y., et al.
Published: (2026)
Towards Efficient Automatic Self-Pruning of Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
From Classification to Generation: An Open-Ended Paradigm for Adverse Drug Reaction Prediction Based on Graph-Motif Feature Fusion
by: Pi, Yuyan, et al.
Published: (2026)
by: Pi, Yuyan, et al.
Published: (2026)
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
OMPQ: Orthogonal Mixed Precision Quantization
by: Ma, Yuexiao, et al.
Published: (2021)
by: Ma, Yuexiao, et al.
Published: (2021)
AffineQuant: Affine Transformation Quantization for Large Language Models
by: Ma, Yuexiao, et al.
Published: (2024)
by: Ma, Yuexiao, et al.
Published: (2024)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
by: Yang, Kaisen, et al.
Published: (2025)
by: Yang, Kaisen, et al.
Published: (2025)
RLPR: Extrapolating RLVR to General Domains without Verifiers
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
Generative AI Models for Different Steps in Architectural Design: A Literature Review
by: Li, Chengyuan, et al.
Published: (2024)
by: Li, Chengyuan, et al.
Published: (2024)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
by: Cheng, Yu, et al.
Published: (2026)
by: Cheng, Yu, et al.
Published: (2026)
Formally Verifying Analog Neural Networks Under Process Variations Using Polynomial Zonotopes
by: Abu-Haeyeh, Yasmine, et al.
Published: (2026)
by: Abu-Haeyeh, Yasmine, et al.
Published: (2026)
Escaping the Verifier: Learning to Reason via Demonstrations
by: Cai, Locke, et al.
Published: (2025)
by: Cai, Locke, et al.
Published: (2025)
Similar Items
-
ALGOGEN: Tool-Generated Verifiable Traces for Reliable Algorithm Visualization
by: Liao, Kunpeng, et al.
Published: (2026) -
Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
by: Ma, Qingchuan, et al.
Published: (2025) -
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
by: Zheng, Xuzhe, et al.
Published: (2026) -
Learning to Generate Formally Verifiable Step-by-Step Logic Reasoning via Structured Formal Intermediaries
by: Chen, Luoxin, et al.
Published: (2026) -
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
by: Xie, Tianyu, et al.
Published: (2026)