On the Power of Decision Trees in Auto-Regressive Language Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gan, Yulu, Galanti, Tomer, Poggio, Tomaso, Malach, Eran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023)
von: Malach, Eran
Veröffentlicht: (2023)
Tool Building as a Path to "Superintelligence"
von: Koplow, David, et al.
Veröffentlicht: (2026)
von: Koplow, David, et al.
Veröffentlicht: (2026)
Unraveling Syntax: How Language Models Learn Context-Free Grammars
von: Schulz, Laura Ying, et al.
Veröffentlicht: (2025)
von: Schulz, Laura Ying, et al.
Veröffentlicht: (2025)
The Fair Language Model Paradox
von: Pinto, Andrea, et al.
Veröffentlicht: (2024)
von: Pinto, Andrea, et al.
Veröffentlicht: (2024)
LLM Priors for ERM over Programs
von: Singhal, Shivam, et al.
Veröffentlicht: (2025)
von: Singhal, Shivam, et al.
Veröffentlicht: (2025)
Formation of Representations in Neural Networks
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
von: Galanti, Tomer, et al.
Veröffentlicht: (2022)
Agentic Systems as Boosting Weak Reasoning Models
von: Sunkaraneni, Varun, et al.
Veröffentlicht: (2026)
von: Sunkaraneni, Varun, et al.
Veröffentlicht: (2026)
Probing Neural Topology of Large Language Models
von: Zheng, Yu, et al.
Veröffentlicht: (2025)
von: Zheng, Yu, et al.
Veröffentlicht: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
PIXAR: Auto-Regressive Language Modeling in Pixel Space
von: Tai, Yintao, et al.
Veröffentlicht: (2024)
von: Tai, Yintao, et al.
Veröffentlicht: (2024)
Self-Assembly of a Biologically Plausible Learning Circuit
von: Liao, Qianli, et al.
Veröffentlicht: (2024)
von: Liao, Qianli, et al.
Veröffentlicht: (2024)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
The Generalized Turing Test: A Foundation for Comparing Intelligence
von: Mitropolsky, Daniel, et al.
Veröffentlicht: (2026)
von: Mitropolsky, Daniel, et al.
Veröffentlicht: (2026)
EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling
von: Ren, Siyu, et al.
Veröffentlicht: (2023)
von: Ren, Siyu, et al.
Veröffentlicht: (2023)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
von: Prabhakar, Akshara, et al.
Veröffentlicht: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
The Role of Sparsity for Length Generalization in Transformers
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
Correlation Dimension of Auto-Regressive Large Language Models
von: Du, Xin, et al.
Veröffentlicht: (2025)
von: Du, Xin, et al.
Veröffentlicht: (2025)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
von: Chijiwa, Daiki, et al.
Veröffentlicht: (2025)
Efficient Context Propagating Perceiver Architectures for Auto-Regressive Language Modeling
von: Mahmood, Kaleel, et al.
Veröffentlicht: (2024)
von: Mahmood, Kaleel, et al.
Veröffentlicht: (2024)
Stop-Think-AutoRegress: Language Modeling with Latent Diffusion Planning
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding
von: Liu, Mingdao, et al.
Veröffentlicht: (2024)
von: Liu, Mingdao, et al.
Veröffentlicht: (2024)
Relation Also Knows: Rethinking the Recall and Editing of Factual Associations in Auto-Regressive Transformer Language Models
von: Liu, Xiyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiyu, et al.
Veröffentlicht: (2024)
On efficiently computable functions, deep networks and sparse compositionality
von: Poggio, Tomaso
Veröffentlicht: (2025)
von: Poggio, Tomaso
Veröffentlicht: (2025)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
von: Ullman, Tomer
Veröffentlicht: (2024)
von: Ullman, Tomer
Veröffentlicht: (2024)
Training the Untrainable: Introducing Inductive Bias via Representational Alignment
von: Subramaniam, Vighnesh, et al.
Veröffentlicht: (2024)
von: Subramaniam, Vighnesh, et al.
Veröffentlicht: (2024)
What's the Plan? Evaluating and Developing Planning-Aware Techniques for Language Models
von: Hirsch, Eran, et al.
Veröffentlicht: (2024)
von: Hirsch, Eran, et al.
Veröffentlicht: (2024)
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Annotations Mitigate Post-Training Mode Collapse
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2026)
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2026)
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
von: Huang, Allison, et al.
Veröffentlicht: (2024)
von: Huang, Allison, et al.
Veröffentlicht: (2024)
A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models
von: Linder, Noa, et al.
Veröffentlicht: (2026)
von: Linder, Noa, et al.
Veröffentlicht: (2026)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2023)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
von: Qin, Tian, et al.
Veröffentlicht: (2025)
von: Qin, Tian, et al.
Veröffentlicht: (2025)
SeaAlert: Critical Information Extraction From Maritime Distress Communications with Large Language Models
von: Atia, Tomer, et al.
Veröffentlicht: (2026)
von: Atia, Tomer, et al.
Veröffentlicht: (2026)
Making Retrieval-Augmented Language Models Robust to Irrelevant Context
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
von: Yoran, Ori, et al.
Veröffentlicht: (2023)
AutoMix: Automatically Mixing Language Models
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2023)
AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
von: Luo, Feng, et al.
Veröffentlicht: (2025)
von: Luo, Feng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Auto-Regressive Next-Token Predictors are Universal Learners
von: Malach, Eran
Veröffentlicht: (2023) -
Tool Building as a Path to "Superintelligence"
von: Koplow, David, et al.
Veröffentlicht: (2026) -
Unraveling Syntax: How Language Models Learn Context-Free Grammars
von: Schulz, Laura Ying, et al.
Veröffentlicht: (2025) -
The Fair Language Model Paradox
von: Pinto, Andrea, et al.
Veröffentlicht: (2024) -
LLM Priors for ERM over Programs
von: Singhal, Shivam, et al.
Veröffentlicht: (2025)