REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Zhuoshi, Pei, Qizhi, Li, Yu, Sun, Qiyao, Tang, Zinan, Zhao, H. Vicky, He, Conghui, Wu, Lijun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
by: Pei, Qizhi, et al.
Published: (2025)
by: Pei, Qizhi, et al.
Published: (2025)
LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
by: Pan, Zhuoshi, et al.
Published: (2025)
by: Pan, Zhuoshi, et al.
Published: (2025)
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
by: Tang, Zinan, et al.
Published: (2025)
by: Tang, Zinan, et al.
Published: (2025)
MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer
by: Lin, Honglin, et al.
Published: (2025)
by: Lin, Honglin, et al.
Published: (2025)
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
by: Lin, Honglin, et al.
Published: (2025)
by: Lin, Honglin, et al.
Published: (2025)
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
by: Pei, Qizhi, et al.
Published: (2025)
by: Pei, Qizhi, et al.
Published: (2025)
A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
by: Ming, Chenlin, et al.
Published: (2025)
by: Ming, Chenlin, et al.
Published: (2025)
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
by: Cai, Mengzhang, et al.
Published: (2025)
by: Cai, Mengzhang, et al.
Published: (2025)
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
A Conformalized Empirical Bayes Method for Multiple Testing with Side Information
by: Zhao, Zinan, et al.
Published: (2025)
by: Zhao, Zinan, et al.
Published: (2025)
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
by: Liu, Zheng, et al.
Published: (2026)
by: Liu, Zheng, et al.
Published: (2026)
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
False Discovery Rate Control For Structured Multiple Testing: Asymmetric Rules And Conformal Q-values
by: Zhao, Zinan, et al.
Published: (2023)
by: Zhao, Zinan, et al.
Published: (2023)
InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior
by: Wang, Huisheng, et al.
Published: (2025)
by: Wang, Huisheng, et al.
Published: (2025)
Conformalized Multiple Testing under Unknown Null Distribution with Symmetric Errors
by: Tian, Yang, et al.
Published: (2025)
by: Tian, Yang, et al.
Published: (2025)
Structure-Adaptive Conformal Inference for Large-Scale Out-of-Distribution Testing
by: Sun, Rongyi, et al.
Published: (2026)
by: Sun, Rongyi, et al.
Published: (2026)
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
by: Lin, Honglin, et al.
Published: (2026)
by: Lin, Honglin, et al.
Published: (2026)
Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models
by: Chen, Danchun, et al.
Published: (2026)
by: Chen, Danchun, et al.
Published: (2026)
FABind+: Enhancing Molecular Docking through Improved Pocket Prediction and Pose Generation
by: Gao, Kaiyuan, et al.
Published: (2024)
by: Gao, Kaiyuan, et al.
Published: (2024)
3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text Modeling
by: Pei, Qizhi, et al.
Published: (2024)
by: Pei, Qizhi, et al.
Published: (2024)
Leveraging Large Language Models to Improve REST API Testing
by: Kim, Myeongsoo, et al.
Published: (2023)
by: Kim, Myeongsoo, et al.
Published: (2023)
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Lightweight and Interpretable Transformer via Mixed Graph Algorithm Unrolling for Traffic Forecast
by: Qi, Ji, et al.
Published: (2025)
by: Qi, Ji, et al.
Published: (2025)
From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion Models
by: Pan, Zhuoshi, et al.
Published: (2023)
by: Pan, Zhuoshi, et al.
Published: (2023)
Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
by: Wei, Jingxuan, et al.
Published: (2025)
by: Wei, Jingxuan, et al.
Published: (2025)
Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
by: Li, Tingyu, et al.
Published: (2025)
by: Li, Tingyu, et al.
Published: (2025)
Log-based, Business-aware REST API Testing
by: Yang, Ding, et al.
Published: (2026)
by: Yang, Ding, et al.
Published: (2026)
DeepREST: Automated Test Case Generation for REST APIs Exploiting Deep Reinforcement Learning
by: Corradini, Davide, et al.
Published: (2024)
by: Corradini, Davide, et al.
Published: (2024)
LRAS: Advanced Legal Reasoning with Agentic Search
by: Zhou, Yujin, et al.
Published: (2026)
by: Zhou, Yujin, et al.
Published: (2026)
An Investigation of Large Language Models and Their Vulnerabilities in Spam Detection
by: Tang, Qiyao, et al.
Published: (2025)
by: Tang, Qiyao, et al.
Published: (2025)
Exploiting Pre-trained Models for Drug Target Affinity Prediction with Nearest Neighbors
by: Pei, Qizhi, et al.
Published: (2024)
by: Pei, Qizhi, et al.
Published: (2024)
Generating REST API Tests With Descriptive Names
by: Garrett, Philip, et al.
Published: (2025)
by: Garrett, Philip, et al.
Published: (2025)
GLEAN: Grounded Lightweight Evaluation Anchors for Contamination-Aware Tabular Reasoning
by: Wang, Qizhi
Published: (2026)
by: Wang, Qizhi
Published: (2026)
Test Amplification for REST APIs Using "Out-of-the-box" Large Language Models
by: Bardakci, Tolgahan, et al.
Published: (2025)
by: Bardakci, Tolgahan, et al.
Published: (2025)
Asking the Right Questions: Improving Reasoning with Generated Stepping Stones
by: Hu, Hengyuan, et al.
Published: (2026)
by: Hu, Hengyuan, et al.
Published: (2026)
You Can REST Now: Automated REST API Documentation and Testing via LLM-Assisted Request Mutations
by: Decrop, Alix, et al.
Published: (2024)
by: Decrop, Alix, et al.
Published: (2024)
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
by: Wei, Jingxuan, et al.
Published: (2025)
by: Wei, Jingxuan, et al.
Published: (2025)
REST: Retrieval-Based Speculative Decoding
by: He, Zhenyu, et al.
Published: (2023)
by: He, Zhenyu, et al.
Published: (2023)
Similar Items
-
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
by: Pei, Qizhi, et al.
Published: (2025) -
LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
by: Pan, Zhuoshi, et al.
Published: (2025) -
Middo: Model-Informed Dynamic Data Optimization for Enhanced LLM Fine-Tuning via Closed-Loop Learning
by: Tang, Zinan, et al.
Published: (2025) -
MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer
by: Lin, Honglin, et al.
Published: (2025) -
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
by: Lin, Honglin, et al.
Published: (2025)