CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Mohtashami, Amirkeivan, Pagliardini, Matteo, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024)
by: Pagliardini, Matteo, et al.
Published: (2024)
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023)
by: Fan, Simin, et al.
Published: (2023)
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025)
by: González, Ramón Calvo, et al.
Published: (2025)
Social Learning: Towards Collaborative Learning with Large Language Models
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025)
by: Semenov, Andrei, et al.
Published: (2025)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
Understanding Hidden Computations in Chain-of-Thought Reasoning
by: Bharadwaj, Aryasomayajula Ram
Published: (2024)
by: Bharadwaj, Aryasomayajula Ram
Published: (2024)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Latent Thought Models with Variational Bayes Inference-Time Computation
by: Kong, Deqian, et al.
Published: (2025)
by: Kong, Deqian, et al.
Published: (2025)
CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
ExpThink: Experience-Guided Reinforcement Learning for Adaptive Chain-of-Thought Compression
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
by: Wagner, Nicolas, et al.
Published: (2024)
by: Wagner, Nicolas, et al.
Published: (2024)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling
by: Xu, Weijia, et al.
Published: (2023)
by: Xu, Weijia, et al.
Published: (2023)
Verifying Chain-of-Thought Reasoning via Its Computational Graph
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference
by: Sabry, Mohammed, et al.
Published: (2026)
by: Sabry, Mohammed, et al.
Published: (2026)
CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity
by: Berchansky, Moshe, et al.
Published: (2024)
by: Berchansky, Moshe, et al.
Published: (2024)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Demystifying Long Chain-of-Thought Reasoning in LLMs
by: Yeo, Edward, et al.
Published: (2025)
by: Yeo, Edward, et al.
Published: (2025)
Constraint-Rectified Training for Efficient Chain-of-Thought
by: Wu, Qinhang, et al.
Published: (2026)
by: Wu, Qinhang, et al.
Published: (2026)
Reliable Chain-of-Thought via Prefix Consistency
by: Iwase, Naoto, et al.
Published: (2026)
by: Iwase, Naoto, et al.
Published: (2026)
M3Hop-CoT: Misogynous Meme Identification with Multimodal Multi-hop Chain-of-Thought
by: Kumari, Gitanjali, et al.
Published: (2024)
by: Kumari, Gitanjali, et al.
Published: (2024)
A Formal Comparison Between Chain of Thought and Latent Thought
by: Xu, Kevin, et al.
Published: (2025)
by: Xu, Kevin, et al.
Published: (2025)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning
by: Zhou, Junjie, et al.
Published: (2025)
by: Zhou, Junjie, et al.
Published: (2025)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
by: Choi, Sehyun
Published: (2024)
by: Choi, Sehyun
Published: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
by: Zhang, Yunfan, et al.
Published: (2025)
by: Zhang, Yunfan, et al.
Published: (2025)
Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning
by: Gu, Yu, et al.
Published: (2026)
by: Gu, Yu, et al.
Published: (2026)
Fractured Chain-of-Thought Reasoning
by: Liao, Baohao, et al.
Published: (2025)
by: Liao, Baohao, et al.
Published: (2025)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
by: Ye, Jiacheng, et al.
Published: (2024)
by: Ye, Jiacheng, et al.
Published: (2024)
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
by: Mu, Yongyu, et al.
Published: (2025)
by: Mu, Yongyu, et al.
Published: (2025)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
by: Kothapalli, Vignesh, et al.
Published: (2025)
by: Kothapalli, Vignesh, et al.
Published: (2025)
Leveraging Large Language Models with Chain-of-Thought and Prompt Engineering for Traffic Crash Severity Analysis and Inference
by: Zhen, Hao, et al.
Published: (2024)
by: Zhen, Hao, et al.
Published: (2024)
Demystifying Chains, Trees, and Graphs of Thoughts
by: Besta, Maciej, et al.
Published: (2024)
by: Besta, Maciej, et al.
Published: (2024)
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024)
by: Bentham, Oliver, et al.
Published: (2024)
Similar Items
-
DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
by: Pagliardini, Matteo, et al.
Published: (2024) -
DoGE: Domain Reweighting with Generalization Estimation
by: Fan, Simin, et al.
Published: (2023) -
Leveraging the true depth of LLMs
by: González, Ramón Calvo, et al.
Published: (2025) -
Social Learning: Towards Collaborative Learning with Large Language Models
by: Mohtashami, Amirkeivan, et al.
Published: (2023) -
Benchmarking Optimizers for Large Language Model Pretraining
by: Semenov, Andrei, et al.
Published: (2025)