Privacy-preserved LLM Cascade via CoT-enhanced Policy Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Kai, Wang, Congchao, Peng, Liqian, Go, Alec, Liu, Xiaozhong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cascade-Aware Training of Language Models
by: Wang, Congchao, et al.
Published: (2024)
by: Wang, Congchao, et al.
Published: (2024)
Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
by: Chen, Wei-Lin, et al.
Published: (2026)
by: Chen, Wei-Lin, et al.
Published: (2026)
Personalized LLM Response Generation with Parameterized Memory Injection
by: Zhang, Kai, et al.
Published: (2024)
by: Zhang, Kai, et al.
Published: (2024)
Syzygy of Thoughts: Improving LLM CoT with the Minimal Free Resolution
by: Li, Chenghao, et al.
Published: (2025)
by: Li, Chenghao, et al.
Published: (2025)
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
by: Jin, Senjie, et al.
Published: (2025)
by: Jin, Senjie, et al.
Published: (2025)
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
You Only Fine-tune Once: Many-Shot In-Context Fine-Tuning for Large Language Models
by: He, Wenchong, et al.
Published: (2025)
by: He, Wenchong, et al.
Published: (2025)
AS-ES Learning: Towards Efficient CoT Learning in Small Models
by: Xi, Nuwa, et al.
Published: (2024)
by: Xi, Nuwa, et al.
Published: (2024)
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination
by: Zhang, Kai, et al.
Published: (2023)
by: Zhang, Kai, et al.
Published: (2023)
ERA-CoT: Improving Chain-of-Thought through Entity Relationship Analysis
by: Liu, Yanming, et al.
Published: (2024)
by: Liu, Yanming, et al.
Published: (2024)
Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation
by: Dai, Chengwei, et al.
Published: (2024)
by: Dai, Chengwei, et al.
Published: (2024)
Can We Trust a Black-box LLM? LLM Untrustworthy Boundary Detection via Bias-Diffusion and Multi-Agent Reinforcement Learning
by: Zhou, Xiaotian, et al.
Published: (2026)
by: Zhou, Xiaotian, et al.
Published: (2026)
Universal Model Routing for Efficient LLM Inference
by: Jitkrittum, Wittawat, et al.
Published: (2025)
by: Jitkrittum, Wittawat, et al.
Published: (2025)
Investigating CoT Monitorability in Large Reasoning Models
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
by: Zheng, Tianshi, et al.
Published: (2025)
by: Zheng, Tianshi, et al.
Published: (2025)
Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
by: Wang, Zezhong, et al.
Published: (2024)
by: Wang, Zezhong, et al.
Published: (2024)
CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
by: Gan, Zeyu, et al.
Published: (2025)
by: Gan, Zeyu, et al.
Published: (2025)
Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
by: Kumarage, Tharindu, et al.
Published: (2025)
by: Kumarage, Tharindu, et al.
Published: (2025)
Investigating Mysteries of CoT-Augmented Distillation
by: Wadhwa, Somin, et al.
Published: (2024)
by: Wadhwa, Somin, et al.
Published: (2024)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
by: Sprague, Zayne, et al.
Published: (2024)
by: Sprague, Zayne, et al.
Published: (2024)
SIM-CoT: Supervised Implicit Chain-of-Thought
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
CoT Vectors: Transferring and Probing the Reasoning Mechanisms of LLMs
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
by: Yan, Shaotian, et al.
Published: (2026)
by: Yan, Shaotian, et al.
Published: (2026)
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think
by: Lee, Seongyun, et al.
Published: (2025)
by: Lee, Seongyun, et al.
Published: (2025)
Efficient Long CoT Reasoning in Small Language Models
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning
by: Wu, Yang, et al.
Published: (2024)
by: Wu, Yang, et al.
Published: (2024)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
by: Li, Jiatong, et al.
Published: (2025)
by: Li, Jiatong, et al.
Published: (2025)
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
by: Ding, Dongyi, et al.
Published: (2025)
by: Ding, Dongyi, et al.
Published: (2025)
Nash CoT: Multi-Path Inference with Preference Equilibrium
by: Zhang, Ziqi, et al.
Published: (2024)
by: Zhang, Ziqi, et al.
Published: (2024)
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
by: Feng, Kehua, et al.
Published: (2025)
by: Feng, Kehua, et al.
Published: (2025)
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
by: Le, Chenqian, et al.
Published: (2025)
by: Le, Chenqian, et al.
Published: (2025)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
by: Luo, Haotian, et al.
Published: (2025)
by: Luo, Haotian, et al.
Published: (2025)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
by: Li, Caihua, et al.
Published: (2024)
by: Li, Caihua, et al.
Published: (2024)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
by: Zhou, Weibo, et al.
Published: (2025)
by: Zhou, Weibo, et al.
Published: (2025)
Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic
by: Zheng, Xin, et al.
Published: (2024)
by: Zheng, Xin, et al.
Published: (2024)
CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
by: Tong, Zhao, et al.
Published: (2026)
by: Tong, Zhao, et al.
Published: (2026)
Generating Effective CoT Traces for Mitigating Causal Hallucination
by: Zhao, Yiheng, et al.
Published: (2026)
by: Zhao, Yiheng, et al.
Published: (2026)
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
by: Wang, Yifei
Published: (2025)
by: Wang, Yifei
Published: (2025)
Similar Items
-
Cascade-Aware Training of Language Models
by: Wang, Congchao, et al.
Published: (2024) -
Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
by: Chen, Wei-Lin, et al.
Published: (2026) -
Personalized LLM Response Generation with Parameterized Memory Injection
by: Zhang, Kai, et al.
Published: (2024) -
Syzygy of Thoughts: Improving LLM CoT with the Minimal Free Resolution
by: Li, Chenghao, et al.
Published: (2025) -
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
by: Jin, Senjie, et al.
Published: (2025)