State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Aviss, Thea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence
von: Aviss, Thea
Veröffentlicht: (2025)
von: Aviss, Thea
Veröffentlicht: (2025)
Improving Embedding Accuracy for Document Retrieval Using Entity Relationship Maps and Model-Aware Contrastive Sampling
von: Aviss, Thea
Veröffentlicht: (2024)
von: Aviss, Thea
Veröffentlicht: (2024)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
von: Huang, Zeyi, et al.
Veröffentlicht: (2026)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
Thinking into the Future: Latent Lookahead Training for Transformers
von: Noci, Lorenzo, et al.
Veröffentlicht: (2026)
von: Noci, Lorenzo, et al.
Veröffentlicht: (2026)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
von: Lu, Wenquan, et al.
Veröffentlicht: (2025)
von: Lu, Wenquan, et al.
Veröffentlicht: (2025)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Parallel Test-Time Scaling for Latent Reasoning Models
von: You, Runyang, et al.
Veröffentlicht: (2025)
von: You, Runyang, et al.
Veröffentlicht: (2025)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
von: Badger, Benjamin L.
Veröffentlicht: (2026)
von: Badger, Benjamin L.
Veröffentlicht: (2026)
Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
von: Liu, Houjun, et al.
Veröffentlicht: (2025)
von: Liu, Houjun, et al.
Veröffentlicht: (2025)
Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence
von: Xiao, Liu
Veröffentlicht: (2026)
von: Xiao, Liu
Veröffentlicht: (2026)
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
von: Su, Guinan, et al.
Veröffentlicht: (2026)
von: Su, Guinan, et al.
Veröffentlicht: (2026)
Reasoning with Latent Thoughts: On the Power of Looped Transformers
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2025)
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2025)
Exploring the Impact of a Transformer's Latent Space Geometry on Downstream Task Performance
von: Marbut, Anna C., et al.
Veröffentlicht: (2024)
von: Marbut, Anna C., et al.
Veröffentlicht: (2024)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
von: Kohli, Harsh, et al.
Veröffentlicht: (2026)
von: Kohli, Harsh, et al.
Veröffentlicht: (2026)
Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
Enhancing Latent Computation in Transformers with Latent Tokens
von: Sun, Yuchang, et al.
Veröffentlicht: (2025)
von: Sun, Yuchang, et al.
Veröffentlicht: (2025)
Geometric Organization of Cognitive States in Transformer Embedding Spaces
von: Zhao, Sophie
Veröffentlicht: (2025)
von: Zhao, Sophie
Veröffentlicht: (2025)
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
von: Jia, Sheng, et al.
Veröffentlicht: (2025)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG
von: Zheng, Yijia, et al.
Veröffentlicht: (2026)
von: Zheng, Yijia, et al.
Veröffentlicht: (2026)
Advancing Regular Language Reasoning in Linear Recurrent Neural Networks
von: Fan, Ting-Han, et al.
Veröffentlicht: (2023)
von: Fan, Ting-Han, et al.
Veröffentlicht: (2023)
Latent Adversarial Training Improves the Representation of Refusal
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures
von: Hoang, Nhat M., et al.
Veröffentlicht: (2025)
von: Hoang, Nhat M., et al.
Veröffentlicht: (2025)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
von: Shen, Shuaijie, et al.
Veröffentlicht: (2024)
von: Shen, Shuaijie, et al.
Veröffentlicht: (2024)
Residual Matrix Transformers: Scaling the Size of the Residual Stream
von: Mak, Brian, et al.
Veröffentlicht: (2025)
von: Mak, Brian, et al.
Veröffentlicht: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
Teaching Transformers Causal Reasoning through Axiomatic Training
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2024)
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2024)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
von: Yang, Songlin, et al.
Veröffentlicht: (2024)
von: Yang, Songlin, et al.
Veröffentlicht: (2024)
Revisiting Bi-Linear State Transitions in Recurrent Neural Networks
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2025)
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2025)
Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V
von: Shinde, Chirag
Veröffentlicht: (2026)
von: Shinde, Chirag
Veröffentlicht: (2026)
Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?
von: Yin, Yutong, et al.
Veröffentlicht: (2025)
von: Yin, Yutong, et al.
Veröffentlicht: (2025)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
von: Li, Hengli, et al.
Veröffentlicht: (2025)
von: Li, Hengli, et al.
Veröffentlicht: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
Inference-Time Rethinking with Latent Thought Vectors for Math Reasoning
von: Kong, Deqian, et al.
Veröffentlicht: (2026)
von: Kong, Deqian, et al.
Veröffentlicht: (2026)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
von: Enström, Daniel, et al.
Veröffentlicht: (2024)
von: Enström, Daniel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
State Stream Transformer (SST) : Emergent Metacognitive Behaviours Through Latent State Persistence
von: Aviss, Thea
Veröffentlicht: (2025) -
Improving Embedding Accuracy for Document Retrieval Using Entity Relationship Maps and Model-Aware Contrastive Sampling
von: Aviss, Thea
Veröffentlicht: (2024) -
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
von: Huang, Zeyi, et al.
Veröffentlicht: (2026) -
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025) -
Thinking into the Future: Latent Lookahead Training for Transformers
von: Noci, Lorenzo, et al.
Veröffentlicht: (2026)