Training Transformers with Enforced Lipschitz Constants
Fuente:
arXiv
Guardado en:
| Autores principales: | Newhouse, Laker, Hess, R. Preston, Cesista, Franz, Zahorodnii, Andrii, Bernstein, Jeremy, Isola, Phillip |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Old Optimizer, New Norm: An Anthology
por: Bernstein, Jeremy, et al.
Publicado: (2024)
por: Bernstein, Jeremy, et al.
Publicado: (2024)
Modular Duality in Deep Learning
por: Bernstein, Jeremy, et al.
Publicado: (2024)
por: Bernstein, Jeremy, et al.
Publicado: (2024)
Improving World Models using Deep Supervision with Linear Probes
por: Zahorodnii, Andrii
Publicado: (2025)
por: Zahorodnii, Andrii
Publicado: (2025)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
por: Huh, Minyoung, et al.
Publicado: (2024)
por: Huh, Minyoung, et al.
Publicado: (2024)
Scalable Optimization in the Modular Norm
por: Large, Tim, et al.
Publicado: (2024)
por: Large, Tim, et al.
Publicado: (2024)
Improved Representation of Asymmetrical Distances with Interval Quasimetric Embeddings
por: Wang, Tongzhou, et al.
Publicado: (2022)
por: Wang, Tongzhou, et al.
Publicado: (2022)
ANTN: Bridging Autoregressive Neural Networks and Tensor Networks for Quantum Many-Body Simulation
por: Chen, Zhuo, et al.
Publicado: (2023)
por: Chen, Zhuo, et al.
Publicado: (2023)
Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
por: Gan, Yulu, et al.
Publicado: (2026)
por: Gan, Yulu, et al.
Publicado: (2026)
Retrieval Augmented Structured Generation: Business Document Information Extraction As Tool Use
por: Cesista, Franz Louis, et al.
Publicado: (2024)
por: Cesista, Franz Louis, et al.
Publicado: (2024)
On the Lipschitz Constant of Deep Networks and Double Descent
por: Gamba, Matteo, et al.
Publicado: (2023)
por: Gamba, Matteo, et al.
Publicado: (2023)
The Truth Lies Somewhere in the Middle (of the Generated Tokens)
por: Wang, Sophie L., et al.
Publicado: (2026)
por: Wang, Sophie L., et al.
Publicado: (2026)
Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report
por: Cesista, Franz Louis
Publicado: (2024)
por: Cesista, Franz Louis
Publicado: (2024)
Chordal Sparsity for Lipschitz Constant Estimation of Deep Neural Networks
por: Xue, Anton, et al.
Publicado: (2022)
por: Xue, Anton, et al.
Publicado: (2022)
Project Jenkins: Turning Monkey Neural Data into Robotic Arm Movement, and Back
por: Zahorodnii, Andrii, et al.
Publicado: (2025)
por: Zahorodnii, Andrii, et al.
Publicado: (2025)
Canonicalizing Multimodal Contrastive Representation Learning
por: Gupta, Sharut, et al.
Publicado: (2026)
por: Gupta, Sharut, et al.
Publicado: (2026)
Words That Make Language Models Perceive
por: Wang, Sophie L., et al.
Publicado: (2025)
por: Wang, Sophie L., et al.
Publicado: (2025)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
por: Bahng, Hyojin, et al.
Publicado: (2025)
por: Bahng, Hyojin, et al.
Publicado: (2025)
Neuroprobe: Evaluating Intracranial Brain Responses to Naturalistic Stimuli
por: Zahorodnii, Andrii, et al.
Publicado: (2025)
por: Zahorodnii, Andrii, et al.
Publicado: (2025)
Softmax Attention with Constant Cost per Token
por: Heinsen, Franz A.
Publicado: (2024)
por: Heinsen, Franz A.
Publicado: (2024)
An Assessment of Model-On-Model Deception
por: Heitkoetter, Julius, et al.
Publicado: (2024)
por: Heitkoetter, Julius, et al.
Publicado: (2024)
Approximation Theory for Lipschitz Continuous Transformers
por: Furuya, Takashi, et al.
Publicado: (2026)
por: Furuya, Takashi, et al.
Publicado: (2026)
An Efficient Global Optimization Algorithm with Adaptive Estimates of the Local Lipschitz Constants
por: D'Agostino, Danny
Publicado: (2022)
por: D'Agostino, Danny
Publicado: (2022)
ExLipBaB: Exact Lipschitz Constant Computation for Piecewise Linear Neural Networks
por: Splittgerber, Tom A.
Publicado: (2026)
por: Splittgerber, Tom A.
Publicado: (2026)
Adaptive Length Image Tokenization via Recurrent Allocation
por: Duggal, Shivam, et al.
Publicado: (2024)
por: Duggal, Shivam, et al.
Publicado: (2024)
Fairness Aware Reward Optimization
por: Choi, Ching Lam, et al.
Publicado: (2026)
por: Choi, Ching Lam, et al.
Publicado: (2026)
Estimating Neural Network Robustness via Lipschitz Constant and Architecture Sensitivity
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
ECLipsE: Efficient Compositional Lipschitz Constant Estimation for Deep Neural Networks
por: Xu, Yuezhu, et al.
Publicado: (2024)
por: Xu, Yuezhu, et al.
Publicado: (2024)
Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
por: Gupta, Sharut, et al.
Publicado: (2025)
por: Gupta, Sharut, et al.
Publicado: (2025)
Personalized Representation from Personalized Generation
por: Sundaram, Shobhita, et al.
Publicado: (2024)
por: Sundaram, Shobhita, et al.
Publicado: (2024)
A Spectral Condition for Feature Learning
por: Yang, Greg, et al.
Publicado: (2023)
por: Yang, Greg, et al.
Publicado: (2023)
The Platonic Representation Hypothesis
por: Huh, Minyoung, et al.
Publicado: (2024)
por: Huh, Minyoung, et al.
Publicado: (2024)
End-to-End Training for Unified Tokenization and Latent Denoising
por: Duggal, Shivam, et al.
Publicado: (2026)
por: Duggal, Shivam, et al.
Publicado: (2026)
Local Lipschitz Constant Computation of ReLU-FNNs: Upper Bound Computation with Exactness Verification
por: Ebihara, Yoshio, et al.
Publicado: (2023)
por: Ebihara, Yoshio, et al.
Publicado: (2023)
Single-pass Adaptive Image Tokenization for Minimum Program Search
por: Duggal, Shivam, et al.
Publicado: (2025)
por: Duggal, Shivam, et al.
Publicado: (2025)
Non-Uniform Memory Sampling in Experience Replay
por: Krutsylo, Andrii
Publicado: (2025)
por: Krutsylo, Andrii
Publicado: (2025)
Kolmogorov Complexity Bounds for LLM Steganography and a Perplexity-Based Detection Proxy
por: Shportko, Andrii
Publicado: (2026)
por: Shportko, Andrii
Publicado: (2026)
Enforcing convex constraints in Graph Neural Networks
por: Rashwan, Ahmed, et al.
Publicado: (2025)
por: Rashwan, Ahmed, et al.
Publicado: (2025)
Lipschitz Constant Meets Condition Number: Learning Robust and Compact Deep Neural Networks
por: Feng, Yangqi, et al.
Publicado: (2025)
por: Feng, Yangqi, et al.
Publicado: (2025)
Step by Step: Adaptive Gradient Descent for Training L-Lipschitz Neural Networks
por: Sung, Kyle, et al.
Publicado: (2025)
por: Sung, Kyle, et al.
Publicado: (2025)
Auditing and Enforcing Conditional Fairness via Optimal Transport
por: Ghassemi, Mohsen, et al.
Publicado: (2024)
por: Ghassemi, Mohsen, et al.
Publicado: (2024)
Ejemplares similares
-
Old Optimizer, New Norm: An Anthology
por: Bernstein, Jeremy, et al.
Publicado: (2024) -
Modular Duality in Deep Learning
por: Bernstein, Jeremy, et al.
Publicado: (2024) -
Improving World Models using Deep Supervision with Linear Probes
por: Zahorodnii, Andrii
Publicado: (2025) -
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
por: Huh, Minyoung, et al.
Publicado: (2024) -
Scalable Optimization in the Modular Norm
por: Large, Tim, et al.
Publicado: (2024)