When Can LLMs Learn to Reason with Weak Supervision?
Fuente:
arXiv
Saved in:
| Main Authors: | Rahman, Salman, Shen, Jingyan, Mordvina, Anna, Palangi, Hamid, Gabriel, Saadia, Izmailov, Pavel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
by: Suvarna, Ashima, et al.
Published: (2026)
by: Suvarna, Ashima, et al.
Published: (2026)
Can LLMs Learn to Reason Robustly under Noisy Supervision?
by: Yang, Shenzhi, et al.
Published: (2026)
by: Yang, Shenzhi, et al.
Published: (2026)
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025)
by: Rahman, Salman, et al.
Published: (2025)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025)
by: Handa, Divij, et al.
Published: (2025)
Removing Sandbagging in LLMs by Training with Weak Supervision
by: Ryd, Emil, et al.
Published: (2026)
by: Ryd, Emil, et al.
Published: (2026)
Reverse Thinking Makes LLMs Stronger Reasoners
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
by: Hwang, Jaedong, et al.
Published: (2025)
by: Hwang, Jaedong, et al.
Published: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
by: Kiyani, Shayan, et al.
Published: (2026)
by: Kiyani, Shayan, et al.
Published: (2026)
Weakly Supervised Learning on Large Graphs
by: Prakash, Aditya
Published: (2025)
by: Prakash, Aditya
Published: (2025)
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
by: Ye, Wen, et al.
Published: (2025)
by: Ye, Wen, et al.
Published: (2025)
Advanced Weakly-Supervised Formula Exploration for Neuro-Symbolic Mathematical Reasoning
by: Wu, Yuxuan, et al.
Published: (2025)
by: Wu, Yuxuan, et al.
Published: (2025)
AdaptThink: Reasoning Models Can Learn When to Think
by: Zhang, Jiajie, et al.
Published: (2025)
by: Zhang, Jiajie, et al.
Published: (2025)
Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
by: Yamin, Khurram, et al.
Published: (2025)
by: Yamin, Khurram, et al.
Published: (2025)
A General Framework for Learning from Weak Supervision
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision?
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models
by: Zhang, Nan, et al.
Published: (2025)
by: Zhang, Nan, et al.
Published: (2025)
A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
by: Rahman, Roussel, et al.
Published: (2025)
by: Rahman, Roussel, et al.
Published: (2025)
2D-OOB: Attributing Data Contribution Through Joint Valuation Framework
by: Sun, Yifan, et al.
Published: (2024)
by: Sun, Yifan, et al.
Published: (2024)
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
by: Tiwari, Sparsh, et al.
Published: (2026)
by: Tiwari, Sparsh, et al.
Published: (2026)
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
by: Goyal, Palash, et al.
Published: (2026)
by: Goyal, Palash, et al.
Published: (2026)
Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
Learning Stable Predictors from Weak Supervision under Distribution Shift
by: Shoeibi, Mehrdad, et al.
Published: (2026)
by: Shoeibi, Mehrdad, et al.
Published: (2026)
Meta-Learn Unimodal Signals with Weak Supervision for Multimodal Sentiment Analysis
by: Mai, Sijie, et al.
Published: (2024)
by: Mai, Sijie, et al.
Published: (2024)
When Can Proxies Improve the Sample Complexity of Preference Learning?
by: Zhu, Yuchen, et al.
Published: (2024)
by: Zhu, Yuchen, et al.
Published: (2024)
Can LLMs Score Medical Diagnoses and Clinical Reasoning as well as Expert Panels?
by: Rouillard, Amy, et al.
Published: (2026)
by: Rouillard, Amy, et al.
Published: (2026)
FastBUS: A Fast Bayesian Framework for Unified Weakly-Supervised Learning
by: Wang, Ziquan, et al.
Published: (2026)
by: Wang, Ziquan, et al.
Published: (2026)
Neuro-symbolic Weak Supervision: Theory and Semantics
by: Upreti, Nijesh, et al.
Published: (2025)
by: Upreti, Nijesh, et al.
Published: (2025)
Lifting Embodied World Models for Planning and Control
by: Wang, Alex N., et al.
Published: (2026)
by: Wang, Alex N., et al.
Published: (2026)
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
by: Cui, Yingqian, et al.
Published: (2026)
by: Cui, Yingqian, et al.
Published: (2026)
A Unified and Stable Risk Minimization Framework for Weakly Supervised Learning with Theoretical Guarantees
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
by: Roy, Shuvendu, et al.
Published: (2025)
by: Roy, Shuvendu, et al.
Published: (2025)
MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
by: Epstein, Elliot L., et al.
Published: (2024)
by: Epstein, Elliot L., et al.
Published: (2024)
Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models
by: Saha, Punyajoy, et al.
Published: (2026)
by: Saha, Punyajoy, et al.
Published: (2026)
Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
by: Cheng, Mingyue, et al.
Published: (2025)
by: Cheng, Mingyue, et al.
Published: (2025)
When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
by: Rehman, Mohammad Abdul, et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Base Models Know How to Reason, Thinking Models Learn When
by: Venhoff, Constantin, et al.
Published: (2025)
by: Venhoff, Constantin, et al.
Published: (2025)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
by: Yao, Xu, et al.
Published: (2026)
by: Yao, Xu, et al.
Published: (2026)
Similar Items
-
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
by: Suvarna, Ashima, et al.
Published: (2026) -
Can LLMs Learn to Reason Robustly under Noisy Supervision?
by: Yang, Shenzhi, et al.
Published: (2026) -
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
by: Rahman, Salman, et al.
Published: (2025) -
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
by: Handa, Divij, et al.
Published: (2025) -
Removing Sandbagging in LLMs by Training with Weak Supervision
by: Ryd, Emil, et al.
Published: (2026)