Training Transformers as a Universal Computer
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Ruize, Yang, Chenxiao, Li, Yanhong, McAllester, David |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
PENCIL: Long Thoughts with Short Memory
by: Yang, Chenxiao, et al.
Published: (2025)
by: Yang, Chenxiao, et al.
Published: (2025)
Context-Efficient Retrieval with Factual Decomposition
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
PFformer: A Position-Free Transformer Variant for Extreme-Adaptive Multivariate Time Series Forecasting
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Supercharging Graph Transformers with Advective Diffusion
by: Wu, Qitian, et al.
Published: (2023)
by: Wu, Qitian, et al.
Published: (2023)
Secure and Efficient Access Control for Computer-Use Agents via Context Space
by: Gong, Haochen, et al.
Published: (2025)
by: Gong, Haochen, et al.
Published: (2025)
ComCLIP: Training-Free Compositional Image and Text Matching
by: Jiang, Kenan, et al.
Published: (2022)
by: Jiang, Kenan, et al.
Published: (2022)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
by: Shen, Guobin, et al.
Published: (2026)
by: Shen, Guobin, et al.
Published: (2026)
Unified Training of Universal Time Series Forecasting Transformers
by: Woo, Gerald, et al.
Published: (2024)
by: Woo, Gerald, et al.
Published: (2024)
Hidden Dynamics of Massive Activations in Transformer Training
by: Gallego-Feliciano, Jorge, et al.
Published: (2025)
by: Gallego-Feliciano, Jorge, et al.
Published: (2025)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Learning Divergence Fields for Shift-Robust Graph Representations
by: Wu, Qitian, et al.
Published: (2024)
by: Wu, Qitian, et al.
Published: (2024)
Teaching an Agent to Sketch One Part at a Time
by: Du, Xiaodan, et al.
Published: (2026)
by: Du, Xiaodan, et al.
Published: (2026)
Compute-Optimal Quantization-Aware Training
by: Dremov, Aleksandr, et al.
Published: (2025)
by: Dremov, Aleksandr, et al.
Published: (2025)
SGFormer: Simplifying and Empowering Transformers for Large-Graph Representations
by: Wu, Qitian, et al.
Published: (2023)
by: Wu, Qitian, et al.
Published: (2023)
Optimizer-Induced Low-Dimensional Drift and Transverse Dynamics in Transformer Training
by: Xu, Yongzhong
Published: (2026)
by: Xu, Yongzhong
Published: (2026)
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
by: Qian, Zekun, et al.
Published: (2026)
by: Qian, Zekun, et al.
Published: (2026)
CI4A: Semantic Component Interfaces for Agents Empowering Web Automation
by: Qiu, Zhi, et al.
Published: (2026)
by: Qiu, Zhi, et al.
Published: (2026)
Scalable Heterogeneous Graph Learning via Heterogeneous-aware Orthogonal Prototype Experts
by: Zhou, Wei, et al.
Published: (2026)
by: Zhou, Wei, et al.
Published: (2026)
Spectral Representation-based Reinforcement Learning
by: Gao, Chenxiao, et al.
Published: (2025)
by: Gao, Chenxiao, et al.
Published: (2025)
Tracking spatial temporal details in ultrasound long video via wavelet analysis and memory bank
by: Zhang, Chenxiao, et al.
Published: (2025)
by: Zhang, Chenxiao, et al.
Published: (2025)
De-conflating Preference and Qualification: Constrained Dual-Perspective Reasoning for Job Recommendation with Large Language Models
by: Kan, Bryce, et al.
Published: (2026)
by: Kan, Bryce, et al.
Published: (2026)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Chunk-Distilled Language Modeling
by: Li, Yanhong, et al.
Published: (2024)
by: Li, Yanhong, et al.
Published: (2024)
Realistic threat perception drives intergroup conflict: A causal, dynamic analysis using generative-agent simulations
by: Abdurahman, Suhaib, et al.
Published: (2025)
by: Abdurahman, Suhaib, et al.
Published: (2025)
OCDB: Revisiting Causal Discovery with a Comprehensive Benchmark and Evaluation Framework
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Training Transformers in Cosine Coefficient Space
by: Bergach, Mohamed Amine
Published: (2026)
by: Bergach, Mohamed Amine
Published: (2026)
Emergent Ordinal Geometry in Transformers Trained on Local Comparisons
by: Singh, Nishit
Published: (2026)
by: Singh, Nishit
Published: (2026)
On the Predictive Power of Representation Dispersion in Language Models
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
OmniDrones: An Efficient and Flexible Platform for Reinforcement Learning in Drone Control
by: Xu, Botian, et al.
Published: (2023)
by: Xu, Botian, et al.
Published: (2023)
AI Agents for Web Testing: A Case Study in the Wild
by: Ye, Naimeng, et al.
Published: (2025)
by: Ye, Naimeng, et al.
Published: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
LP-LM: No Hallucinations in Question Answering with Logic Programming
by: Wu, Katherine, et al.
Published: (2025)
by: Wu, Katherine, et al.
Published: (2025)
To See is Not to Master: Teaching LLMs to Use Private Libraries for Code Generation
by: Zhang, Yitong, et al.
Published: (2026)
by: Zhang, Yitong, et al.
Published: (2026)
Quantization-Aware and Tensor-Compressed Training of Transformers for Natural Language Understanding
by: Yang, Zi, et al.
Published: (2023)
by: Yang, Zi, et al.
Published: (2023)
Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning
by: Zhang, Ruize, et al.
Published: (2025)
by: Zhang, Ruize, et al.
Published: (2025)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
by: Xiao, Jinqi, et al.
Published: (2025)
by: Xiao, Jinqi, et al.
Published: (2025)
MESIA: Understanding and Leveraging Supplementary Nature of Method-level Comments for Automatic Comment Generation
by: Pan, Xinglu, et al.
Published: (2024)
by: Pan, Xinglu, et al.
Published: (2024)
Relabeling Minimal Training Subset to Flip a Prediction
by: Yang, Jinghan, et al.
Published: (2023)
by: Yang, Jinghan, et al.
Published: (2023)
Similar Items
-
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
by: Li, Yanhong, et al.
Published: (2025) -
PENCIL: Long Thoughts with Short Memory
by: Yang, Chenxiao, et al.
Published: (2025) -
Context-Efficient Retrieval with Factual Decomposition
by: Li, Yanhong, et al.
Published: (2025) -
PFformer: A Position-Free Transformer Variant for Extreme-Adaptive Multivariate Time Series Forecasting
by: Li, Yanhong, et al.
Published: (2025) -
Supercharging Graph Transformers with Advective Diffusion
by: Wu, Qitian, et al.
Published: (2023)