Small Models Struggle to Learn from Strong Reasoners
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yuetai, Yue, Xiang, Xu, Zhangchen, Jiang, Fengqing, Niu, Luyao, Lin, Bill Yuchen, Ramasubramanian, Bhaskar, Poovendran, Radha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Temporal Sampling for Forgotten Reasoning in LLMs
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
von: Li, Yuetai, et al.
Veröffentlicht: (2024)
von: Li, Yuetai, et al.
Veröffentlicht: (2024)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Stronger Models are NOT Stronger Teachers for Instruction Tuning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks
von: Jiang, Fengqing, et al.
Veröffentlicht: (2026)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2026)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Polyhedral Instability Governs Regret in Online Learning
von: Li, Yuetai, et al.
Veröffentlicht: (2026)
von: Li, Yuetai, et al.
Veröffentlicht: (2026)
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
von: Feng, Yichen, et al.
Veröffentlicht: (2026)
von: Feng, Yichen, et al.
Veröffentlicht: (2026)
JobBench: Aligning Agent Work With Human Will
von: Li, Yuetai, et al.
Veröffentlicht: (2026)
von: Li, Yuetai, et al.
Veröffentlicht: (2026)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
A Method for Fast Autonomy Transfer in Reinforcement Learning
von: Sahabandu, Dinuka, et al.
Veröffentlicht: (2024)
von: Sahabandu, Dinuka, et al.
Veröffentlicht: (2024)
Double-Dip: Thwarting Label-Only Membership Inference Attacks with Transfer Learning and Randomization
von: Rajabi, Arezoo, et al.
Veröffentlicht: (2024)
von: Rajabi, Arezoo, et al.
Veröffentlicht: (2024)
Fault Tolerant Neural Control Barrier Functions for Robotic Systems under Sensor Faults and Attacks
von: Zhang, Hongchao, et al.
Veröffentlicht: (2024)
von: Zhang, Hongchao, et al.
Veröffentlicht: (2024)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
Modeling and Designing Non-Pharmaceutical Interventions in Epidemics: A Submodular Approach
von: Cheng, Shiyu, et al.
Veröffentlicht: (2024)
von: Cheng, Shiyu, et al.
Veröffentlicht: (2024)
Simulating Environments with Reasoning Models for Agent Training
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
Distributed Safety-Critical Control of Multi-Agent Systems with Time-Varying Communication Topologies
von: Cheng, Shiyu, et al.
Veröffentlicht: (2026)
von: Cheng, Shiyu, et al.
Veröffentlicht: (2026)
Swarm-STL: A Framework for Motion Planning in Large-Scale, Multi-Swarm Systems
von: Cheng, Shiyu, et al.
Veröffentlicht: (2025)
von: Cheng, Shiyu, et al.
Veröffentlicht: (2025)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
von: Sahabandu, Dinuka, et al.
Veröffentlicht: (2024)
von: Sahabandu, Dinuka, et al.
Veröffentlicht: (2024)
Who is Responsible? Explaining Safety Violations in Multi-Agent Cyber-Physical Systems
von: Niu, Luyao, et al.
Veröffentlicht: (2024)
von: Niu, Luyao, et al.
Veröffentlicht: (2024)
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
Long-context LLMs Struggle with Long In-context Learning
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
Reasoning Models Struggle to Control their Chains of Thought
von: Yueh-Han, Chen, et al.
Veröffentlicht: (2026)
von: Yueh-Han, Chen, et al.
Veröffentlicht: (2026)
CANTXSec: A Deterministic Intrusion Detection and Prevention System for CAN Bus Monitoring ECU Activations
von: Donadel, Denis, et al.
Veröffentlicht: (2025)
von: Donadel, Denis, et al.
Veröffentlicht: (2025)
Collaborative Agent Reasoning Engineering (CARE): A Three-Party Design Methodology for Systematically Engineering AI Agents with Subject Matter Experts, Developers, and Helper Agents
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2026)
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2026)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
von: Yang, Wang, et al.
Veröffentlicht: (2025)
von: Yang, Wang, et al.
Veröffentlicht: (2025)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
von: Song, Yifan, et al.
Veröffentlicht: (2024)
von: Song, Yifan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Temporal Sampling for Forgotten Reasoning in LLMs
von: Li, Yuetai, et al.
Veröffentlicht: (2025) -
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025) -
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
von: Li, Yuetai, et al.
Veröffentlicht: (2024) -
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
von: Feng, Yichen, et al.
Veröffentlicht: (2025) -
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)