A One-Layer Decoder-Only Transformer is a Two-Layer RNN: With an Application to Certified Robustness
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yuhao, Albarghouthi, Aws, D'Antoni, Loris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
por: Zhang, Yuhao, et al.
Publicado: (2023)
por: Zhang, Yuhao, et al.
Publicado: (2023)
Verified Training for Counterfactual Explanation Robustness under Data Shift
por: Meyer, Anna P., et al.
Publicado: (2024)
por: Meyer, Anna P., et al.
Publicado: (2024)
Perceptions of the Fairness Impacts of Multiplicity in Machine Learning
por: Meyer, Anna P., et al.
Publicado: (2024)
por: Meyer, Anna P., et al.
Publicado: (2024)
Flexible and Efficient Grammar-Constrained Decoding
por: Park, Kanghee, et al.
Publicado: (2025)
por: Park, Kanghee, et al.
Publicado: (2025)
Linear-Time T-Gate Optimization via Random Abstraction
por: Albarghouthi, Aws
Publicado: (2026)
por: Albarghouthi, Aws
Publicado: (2026)
Analyzing Decoders for Quantum Error Correction
por: Molavi, Abtin, et al.
Publicado: (2026)
por: Molavi, Abtin, et al.
Publicado: (2026)
Unrealizability Logic
por: Kim, Jinwoo, et al.
Publicado: (2022)
por: Kim, Jinwoo, et al.
Publicado: (2022)
Grammar-Aligned Decoding
por: Park, Kanghee, et al.
Publicado: (2024)
por: Park, Kanghee, et al.
Publicado: (2024)
The Format Tax
por: Lee, Ivan Yee, et al.
Publicado: (2026)
por: Lee, Ivan Yee, et al.
Publicado: (2026)
Verifying Solutions to Semantics-Guided Synthesis Problems
por: Murphy, Charlie, et al.
Publicado: (2024)
por: Murphy, Charlie, et al.
Publicado: (2024)
LOUD: Synthesizing Strongest and Weakest Specifications
por: Park, Kanghee, et al.
Publicado: (2024)
por: Park, Kanghee, et al.
Publicado: (2024)
Synthesizing Specifications
por: Park, Kanghee, et al.
Publicado: (2023)
por: Park, Kanghee, et al.
Publicado: (2023)
Constrained Adaptive Rejection Sampling
por: Parys, Paweł, et al.
Publicado: (2025)
por: Parys, Paweł, et al.
Publicado: (2025)
Language-Based Agent Control
por: Zhou, Timothy, et al.
Publicado: (2026)
por: Zhou, Timothy, et al.
Publicado: (2026)
Bootstrapping Fuzzers for Compilers of Low-Resource Language Dialects Using Language Models
por: Vaidya, Sairam, et al.
Publicado: (2025)
por: Vaidya, Sairam, et al.
Publicado: (2025)
Semantics of Sets of Programs
por: Kim, Jinwoo, et al.
Publicado: (2024)
por: Kim, Jinwoo, et al.
Publicado: (2024)
Automating Unrealizability Logic: Hoare-Style Proof Synthesis for Infinite Sets of Programs
por: Nagy, Shaan, et al.
Publicado: (2024)
por: Nagy, Shaan, et al.
Publicado: (2024)
Continuous Diffusion Models Can Obey Formal Syntax
por: Kim, Jinwoo, et al.
Publicado: (2026)
por: Kim, Jinwoo, et al.
Publicado: (2026)
ChopChop: a Programmable Framework for Semantically Constraining the Output of Language Models
por: Nagy, Shaan, et al.
Publicado: (2025)
por: Nagy, Shaan, et al.
Publicado: (2025)
Nice to Meet You: Synthesizing Practical MLIR Abstract Transformers
por: Peng, Xuanyu, et al.
Publicado: (2025)
por: Peng, Xuanyu, et al.
Publicado: (2025)
Optimizing Quantum Circuits, Fast and Slow
por: Xu, Amanda, et al.
Publicado: (2024)
por: Xu, Amanda, et al.
Publicado: (2024)
Dependency-Aware Compilation for Surface Code Quantum Architectures
por: Molavi, Abtin, et al.
Publicado: (2023)
por: Molavi, Abtin, et al.
Publicado: (2023)
Pareto Optimal Code Generation
por: Orlanski, Gabriel, et al.
Publicado: (2025)
por: Orlanski, Gabriel, et al.
Publicado: (2025)
The SemGuS Toolkit
por: Johnson, Keith J. C., et al.
Publicado: (2024)
por: Johnson, Keith J. C., et al.
Publicado: (2024)
Automating Pruning in Top-Down Enumeration for Program Synthesis Problems with Monotonic Semantics
por: Johnson, Keith J. C., et al.
Publicado: (2024)
por: Johnson, Keith J. C., et al.
Publicado: (2024)
LayerNorm Induces Recency Bias in Transformer Decoders
por: Kim, Junu, et al.
Publicado: (2025)
por: Kim, Junu, et al.
Publicado: (2025)
Constrained Sampling for Language Models Should Be Easy: An MCMC Perspective
por: Gonzalez, Emmanuel Anaya, et al.
Publicado: (2025)
por: Gonzalez, Emmanuel Anaya, et al.
Publicado: (2025)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
por: Ahrabian, Kian, et al.
Publicado: (2024)
por: Ahrabian, Kian, et al.
Publicado: (2024)
Rethinking the adaptive relationship between Encoder Layers and Decoder Layers
por: Song, Yubo
Publicado: (2024)
por: Song, Yubo
Publicado: (2024)
Provable Knowledge Acquisition and Extraction in One-Layer Transformers
por: Xu, Ruichen, et al.
Publicado: (2025)
por: Xu, Ruichen, et al.
Publicado: (2025)
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
por: Wei, Zhepei, et al.
Publicado: (2025)
por: Wei, Zhepei, et al.
Publicado: (2025)
Generating Compilers for Qubit Mapping and Routing
por: Molavi, Abtin, et al.
Publicado: (2025)
por: Molavi, Abtin, et al.
Publicado: (2025)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
por: Kamigaito, Hidetaka, et al.
Publicado: (2025)
por: Kamigaito, Hidetaka, et al.
Publicado: (2025)
Transformer Layers as Painters
por: Sun, Qi, et al.
Publicado: (2024)
por: Sun, Qi, et al.
Publicado: (2024)
How Powerful are Decoder-Only Transformer Neural Models?
por: Roberts, Jesse
Publicado: (2023)
por: Roberts, Jesse
Publicado: (2023)
Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers
por: He, Zicong, et al.
Publicado: (2025)
por: He, Zicong, et al.
Publicado: (2025)
Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
por: Zhang, Biao, et al.
Publicado: (2025)
por: Zhang, Biao, et al.
Publicado: (2025)
ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding
por: Amer, Walaa, et al.
Publicado: (2026)
por: Amer, Walaa, et al.
Publicado: (2026)
Lower Layers Matter: Alleviating Hallucination via Multi-Layer Fusion Contrastive Decoding with Truthfulness Refocused
por: Chen, Dingwei, et al.
Publicado: (2024)
por: Chen, Dingwei, et al.
Publicado: (2024)
You Only Cache Once: Decoder-Decoder Architectures for Language Models
por: Sun, Yutao, et al.
Publicado: (2024)
por: Sun, Yutao, et al.
Publicado: (2024)
Ejemplares similares
-
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
por: Zhang, Yuhao, et al.
Publicado: (2023) -
Verified Training for Counterfactual Explanation Robustness under Data Shift
por: Meyer, Anna P., et al.
Publicado: (2024) -
Perceptions of the Fairness Impacts of Multiplicity in Machine Learning
por: Meyer, Anna P., et al.
Publicado: (2024) -
Flexible and Efficient Grammar-Constrained Decoding
por: Park, Kanghee, et al.
Publicado: (2025) -
Linear-Time T-Gate Optimization via Random Abstraction
por: Albarghouthi, Aws
Publicado: (2026)