Stochastic activations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lomeli, Maria, Douze, Matthijs, Szilvasy, Gergely, Cabannes, Loic, Copet, Jade, Sukhbaatar, Sainbayar, Weston, Jason, Synnaeve, Gabriel, Mazaré, Pierre-Emmanuel, Jégou, Hervé |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Short window attention enables long-term memorization
von: Cabannes, Loïc, et al.
Veröffentlicht: (2025)
von: Cabannes, Loïc, et al.
Veröffentlicht: (2025)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026)
Vector search with small radiuses
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2024)
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2024)
Inference-time sparse attention with asymmetric indexing
von: Mazaré, Pierre-Emmanuel, et al.
Veröffentlicht: (2025)
von: Mazaré, Pierre-Emmanuel, et al.
Veröffentlicht: (2025)
The Faiss library
von: Douze, Matthijs, et al.
Veröffentlicht: (2024)
von: Douze, Matthijs, et al.
Veröffentlicht: (2024)
Contextual Position Encoding: Learning to Count What's Important
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
von: Xu, Jing, et al.
Veröffentlicht: (2023)
von: Xu, Jing, et al.
Veröffentlicht: (2023)
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
Thinking LLMs: General Instruction Following with Thought Generation
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Iterative Reasoning Preference Optimization
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
von: Pang, Richard Yuanzhe, et al.
Veröffentlicht: (2024)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
StepWiser: Stepwise Generative Judges for Wiser Reasoning
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
R.I.P.: Better Models by Survival of the Fittest Prompts
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Self-Rewarding Language Models
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
von: Yuan, Weizhe, et al.
Veröffentlicht: (2024)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Functional Invariants to Watermark Large Transformers
von: Fernandez, Pierre, et al.
Veröffentlicht: (2023)
von: Fernandez, Pierre, et al.
Veröffentlicht: (2023)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
von: Yu, Ping, et al.
Veröffentlicht: (2025)
von: Yu, Ping, et al.
Veröffentlicht: (2025)
Simple and Controllable Music Generation
von: Copet, Jade, et al.
Veröffentlicht: (2023)
von: Copet, Jade, et al.
Veröffentlicht: (2023)
Automatic Textbook Formalization
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2026)
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2026)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
Masked Audio Generation using a Single Non-Autoregressive Transformer
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
Watermarking Makes Language Models Radioactive
von: Sander, Tom, et al.
Veröffentlicht: (2024)
von: Sander, Tom, et al.
Veröffentlicht: (2024)
Multi-Token Attention
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
von: Golovneva, Olga, et al.
Veröffentlicht: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces
von: Su, DiJia, et al.
Veröffentlicht: (2024)
von: Su, DiJia, et al.
Veröffentlicht: (2024)
Moshi: a speech-text foundation model for real-time dialogue
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping
von: Lehnert, Lucas, et al.
Veröffentlicht: (2024)
von: Lehnert, Lucas, et al.
Veröffentlicht: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
von: Tan, Ellen Xiaoqing, et al.
Veröffentlicht: (2026)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
Neutral Residues: Revisiting Adapters for Model Extension
von: Talla, Franck Signe, et al.
Veröffentlicht: (2024)
von: Talla, Franck Signe, et al.
Veröffentlicht: (2024)
Machine learning and high dimensional vector search
von: Douze, Matthijs
Veröffentlicht: (2025)
von: Douze, Matthijs
Veröffentlicht: (2025)
AI & Human Co-Improvement for Safer Co-Superintelligence
von: Weston, Jason, et al.
Veröffentlicht: (2025)
von: Weston, Jason, et al.
Veröffentlicht: (2025)
RA-DIT: Retrieval-Augmented Dual Instruction Tuning
von: Lin, Xi Victoria, et al.
Veröffentlicht: (2023)
von: Lin, Xi Victoria, et al.
Veröffentlicht: (2023)
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2025)
An Independence-promoting Loss for Music Generation with Language Models
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
Touring sampling with pushforward maps
von: Cabannes, Vivien, et al.
Veröffentlicht: (2023)
von: Cabannes, Vivien, et al.
Veröffentlicht: (2023)
The Galerkin method beats Graph-Based Approaches for Spectral Algorithms
von: Cabannes, Vivien, et al.
Veröffentlicht: (2023)
von: Cabannes, Vivien, et al.
Veröffentlicht: (2023)
In-context Pretraining: Language Modeling Beyond Document Boundaries
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Short window attention enables long-term memorization
von: Cabannes, Loïc, et al.
Veröffentlicht: (2025) -
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2026) -
Vector search with small radiuses
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2024) -
Inference-time sparse attention with asymmetric indexing
von: Mazaré, Pierre-Emmanuel, et al.
Veröffentlicht: (2025) -
The Faiss library
von: Douze, Matthijs, et al.
Veröffentlicht: (2024)