MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Sprague, Zayne, Ye, Xi, Bostrom, Kaj, Chaudhuri, Swarat, Durrett, Greg |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025)
by: Park, Chanwoo, et al.
Published: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024)
by: Sprague, Zayne, et al.
Published: (2024)
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
by: Sprague, Zayne, et al.
Published: (2025)
by: Sprague, Zayne, et al.
Published: (2025)
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025)
by: Wadhwa, Manya, et al.
Published: (2025)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
LoFiT: Localized Fine-tuning on LLM Representations
by: Yin, Fangcong, et al.
Published: (2024)
by: Yin, Fangcong, et al.
Published: (2024)
MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
by: Dutta, Aritra, et al.
Published: (2026)
by: Dutta, Aritra, et al.
Published: (2026)
Learning Composable Chains-of-Thought
by: Yin, Fangcong, et al.
Published: (2025)
by: Yin, Fangcong, et al.
Published: (2025)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
by: Tang, Liyan, et al.
Published: (2025)
by: Tang, Liyan, et al.
Published: (2025)
Batched Low-Rank Adaptation of Foundation Models
by: Wen, Yeming, et al.
Published: (2023)
by: Wen, Yeming, et al.
Published: (2023)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
by: Xin, Yutong, et al.
Published: (2026)
by: Xin, Yutong, et al.
Published: (2026)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
by: Gunjal, Anisha, et al.
Published: (2024)
by: Gunjal, Anisha, et al.
Published: (2024)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
CLEVER: A Curated Benchmark for Formally Verified Code Generation
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
SoftCoT++: Test-Time Scaling with Soft Chain-of-Thought Reasoning
by: Xu, Yige, et al.
Published: (2025)
by: Xu, Yige, et al.
Published: (2025)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
by: Liu, Zeyu Leo, et al.
Published: (2024)
by: Liu, Zeyu Leo, et al.
Published: (2024)
Understanding Synthetic Context Extension via Retrieval Heads
by: Zhao, Xinyu, et al.
Published: (2024)
by: Zhao, Xinyu, et al.
Published: (2024)
CFLOBDDs: Context-Free-Language Ordered Binary Decision Diagrams
by: Sistla, Meghana, et al.
Published: (2022)
by: Sistla, Meghana, et al.
Published: (2022)
CREATE: Testing LLMs for Associative Creativity
by: Wadhwa, Manya, et al.
Published: (2026)
by: Wadhwa, Manya, et al.
Published: (2026)
Learning Quantitative Automata Modulo Theories
by: Hsiung, Eric, et al.
Published: (2024)
by: Hsiung, Eric, et al.
Published: (2024)
X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs
by: Rodriguez, Juan Diego, et al.
Published: (2023)
by: Rodriguez, Juan Diego, et al.
Published: (2023)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
by: Ding, Wenxuan, et al.
Published: (2026)
by: Ding, Wenxuan, et al.
Published: (2026)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
by: Lake, Thom, et al.
Published: (2024)
by: Lake, Thom, et al.
Published: (2024)
Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling
by: Sanders, Kate, et al.
Published: (2026)
by: Sanders, Kate, et al.
Published: (2026)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
by: Ye, Xi, et al.
Published: (2025)
by: Ye, Xi, et al.
Published: (2025)
Using Natural Language Explanations to Rescale Human Judgments
by: Wadhwa, Manya, et al.
Published: (2023)
by: Wadhwa, Manya, et al.
Published: (2023)
RankAlign: A Ranking View of the Generator-Validator Gap in Large Language Models
by: Rodriguez, Juan Diego, et al.
Published: (2025)
by: Rodriguez, Juan Diego, et al.
Published: (2025)
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024)
by: Wadhwa, Manya, et al.
Published: (2024)
Online Cascade Learning for Efficient Inference over Streams
by: Nie, Lunyiu, et al.
Published: (2024)
by: Nie, Lunyiu, et al.
Published: (2024)
SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs
by: Xu, Yige, et al.
Published: (2025)
by: Xu, Yige, et al.
Published: (2025)
A Long Way to Go: Investigating Length Correlations in RLHF
by: Singhal, Prasann, et al.
Published: (2023)
by: Singhal, Prasann, et al.
Published: (2023)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
by: Liu, Zeyu Leo, et al.
Published: (2025)
by: Liu, Zeyu Leo, et al.
Published: (2025)
Complex Claim Verification with Evidence Retrieved in the Wild
by: Chen, Jifan, et al.
Published: (2023)
by: Chen, Jifan, et al.
Published: (2023)
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
by: Jiang, Yuru, et al.
Published: (2025)
by: Jiang, Yuru, et al.
Published: (2025)
D2PO: Discriminator-Guided DPO with Response Evaluation Models
by: Singhal, Prasann, et al.
Published: (2024)
by: Singhal, Prasann, et al.
Published: (2024)
Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
by: Govindarajan, Venkata S, et al.
Published: (2023)
by: Govindarajan, Venkata S, et al.
Published: (2023)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
by: Wang, Songtao, et al.
Published: (2026)
by: Wang, Songtao, et al.
Published: (2026)
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?
by: Zhou, Zhanke, et al.
Published: (2024)
by: Zhou, Zhanke, et al.
Published: (2024)
Early Stopping Chain-of-thoughts in Large Language Models
by: Mao, Minjia, et al.
Published: (2025)
by: Mao, Minjia, et al.
Published: (2025)
Similar Items
-
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
by: Park, Chanwoo, et al.
Published: (2025) -
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024) -
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
by: Sprague, Zayne, et al.
Published: (2025) -
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
by: Wadhwa, Manya, et al.
Published: (2025) -
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
by: Thakur, Amitayush, et al.
Published: (2025)