Mechanistic Design and Scaling of Hybrid Architectures
Fuente:
arXiv
Guardado en:
| Autores principales: | Poli, Michael, Thomas, Armin W, Nguyen, Eric, Ponnusamy, Pragaash, Deiseroth, Björn, Kersting, Kristian, Suzuki, Taiji, Hie, Brian, Ermon, Stefano, Ré, Christopher, Zhang, Ce, Massaroli, Stefano |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Systems and Algorithms for Convolutional Multi-Hybrid Language Models at Scale
por: Ku, Jerome, et al.
Publicado: (2025)
por: Ku, Jerome, et al.
Publicado: (2025)
Sliding Window Recurrences for Sequence Models
por: Secrieru, Dragos, et al.
Publicado: (2025)
por: Secrieru, Dragos, et al.
Publicado: (2025)
STAR: Synthesis of Tailored Architectures
por: Thomas, Armin W., et al.
Publicado: (2024)
por: Thomas, Armin W., et al.
Publicado: (2024)
State-Free Inference of State-Space Models: The Transfer Function Approach
por: Parnichkun, Rom N., et al.
Publicado: (2024)
por: Parnichkun, Rom N., et al.
Publicado: (2024)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
por: Deiseroth, Björn, et al.
Publicado: (2025)
por: Deiseroth, Björn, et al.
Publicado: (2025)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
por: Höth, Max Henning, et al.
Publicado: (2026)
por: Höth, Max Henning, et al.
Publicado: (2026)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
por: Deiseroth, Björn, et al.
Publicado: (2024)
por: Deiseroth, Björn, et al.
Publicado: (2024)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
por: Sztwiertnia, Sebastian, et al.
Publicado: (2025)
por: Sztwiertnia, Sebastian, et al.
Publicado: (2025)
Quantifying Memory Utilization with Effective State-Size
por: Parnichkun, Rom N., et al.
Publicado: (2025)
por: Parnichkun, Rom N., et al.
Publicado: (2025)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
por: Deiseroth, Björn, et al.
Publicado: (2023)
por: Deiseroth, Björn, et al.
Publicado: (2023)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
por: Härle, Ruben, et al.
Publicado: (2024)
por: Härle, Ruben, et al.
Publicado: (2024)
Measuring and Guiding Monosemanticity
por: Härle, Ruben, et al.
Publicado: (2025)
por: Härle, Ruben, et al.
Publicado: (2025)
Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization
por: Deiseroth, Björn, et al.
Publicado: (2023)
por: Deiseroth, Björn, et al.
Publicado: (2023)
Exploring Diffusion Transformer Designs via Grafting
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2025)
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2025)
Training-Free Activation Sparsity in Large Language Models
por: Liu, James, et al.
Publicado: (2024)
por: Liu, James, et al.
Publicado: (2024)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
por: Guo, Gabe, et al.
Publicado: (2025)
por: Guo, Gabe, et al.
Publicado: (2025)
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
por: Cundy, Chris, et al.
Publicado: (2023)
por: Cundy, Chris, et al.
Publicado: (2023)
RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
por: Chen, Tianlang, et al.
Publicado: (2025)
por: Chen, Tianlang, et al.
Publicado: (2025)
Educação Permanente para o aperfeiçoamento do Controle de Infecção Hospitalar: revisão integrativa
por: Aline Massaroli
Publicado: (2014)
por: Aline Massaroli
Publicado: (2014)
Cosmological horizons
por: Fiorentin, Michele Re, et al.
Publicado: (2025)
por: Fiorentin, Michele Re, et al.
Publicado: (2025)
Privacy-Constrained Policies via Mutual Information Regularized Policy Gradients
por: Cundy, Chris, et al.
Publicado: (2020)
por: Cundy, Chris, et al.
Publicado: (2020)
Inductive Moment Matching
por: Zhou, Linqi, et al.
Publicado: (2025)
por: Zhou, Linqi, et al.
Publicado: (2025)
Calibrated Probabilistic Forecasts for Arbitrary Sequences
por: Marx, Charles, et al.
Publicado: (2024)
por: Marx, Charles, et al.
Publicado: (2024)
TrAct: Making First-layer Pre-Activations Trainable
por: Petersen, Felix, et al.
Publicado: (2024)
por: Petersen, Felix, et al.
Publicado: (2024)
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
por: Lou, Aaron, et al.
Publicado: (2023)
por: Lou, Aaron, et al.
Publicado: (2023)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
por: Manvi, Rohin, et al.
Publicado: (2024)
por: Manvi, Rohin, et al.
Publicado: (2024)
Mission Design for Unmanned Aerial Vehicles using Hybrid Probabilistic Logic Programs
por: Kohaut, Simon, et al.
Publicado: (2024)
por: Kohaut, Simon, et al.
Publicado: (2024)
Who's at Risk? Effects of Inflation on Unemployment Risk
por: Ahn, Hie Joo, et al.
Publicado: (2025)
por: Ahn, Hie Joo, et al.
Publicado: (2025)
DistillKac: Few-Step Image Generation via Damped Wave Equations
por: Han, Weiqiao, et al.
Publicado: (2025)
por: Han, Weiqiao, et al.
Publicado: (2025)
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
por: Neitemeier, Pit, et al.
Publicado: (2025)
por: Neitemeier, Pit, et al.
Publicado: (2025)
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
por: Mitchell, Rupert, et al.
Publicado: (2025)
por: Mitchell, Rupert, et al.
Publicado: (2025)
Competitividade e gap tecnológico – uma análise comparativa entre Brasil e países europeus
por: Tatiana Massaroli Melo
Publicado: (2017)
por: Tatiana Massaroli Melo
Publicado: (2017)
Um modelo setorial baseado na abordagem kaleckiana da distribuição setorial funcional da renda e na teoria schumpeteriana da concorrência
por: Tatiana Massaroli Melo
Publicado: (2016)
por: Tatiana Massaroli Melo
Publicado: (2016)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
por: Rafailov, Rafael, et al.
Publicado: (2023)
por: Rafailov, Rafael, et al.
Publicado: (2023)
An energy landscape-based theoretical framework for understanding the emergence of functions in a living system under the dynamical component interaction
por: Suzuki, Ryunosuke, et al.
Publicado: (2025)
por: Suzuki, Ryunosuke, et al.
Publicado: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
por: Higuchi, Rei, et al.
Publicado: (2025)
por: Higuchi, Rei, et al.
Publicado: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
por: Awano, Ryoya, et al.
Publicado: (2026)
por: Awano, Ryoya, et al.
Publicado: (2026)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
por: Nishikawa, Naoki, et al.
Publicado: (2024)
por: Nishikawa, Naoki, et al.
Publicado: (2024)
Ejemplares similares
-
Systems and Algorithms for Convolutional Multi-Hybrid Language Models at Scale
por: Ku, Jerome, et al.
Publicado: (2025) -
Sliding Window Recurrences for Sequence Models
por: Secrieru, Dragos, et al.
Publicado: (2025) -
STAR: Synthesis of Tailored Architectures
por: Thomas, Armin W., et al.
Publicado: (2024) -
State-Free Inference of State-Space Models: The Transfer Function Approach
por: Parnichkun, Rom N., et al.
Publicado: (2024) -
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
por: Deiseroth, Björn, et al.
Publicado: (2025)