Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
Fuente:
arXiv
Salvato in:
| Autori principali: | Potapczynski, Andres, Qiu, Shikai, Finzi, Marc, Ferri, Christopher, Chen, Zixi, Goldblum, Micah, Bruss, Bayan, De Sa, Christopher, Wilson, Andrew Gordon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Compute Better Spent: Replacing Dense Layers with Structured Matrices
di: Qiu, Shikai, et al.
Pubblicazione: (2024)
di: Qiu, Shikai, et al.
Pubblicazione: (2024)
The Lie Derivative for Measuring Learned Equivariance
di: Gruver, Nate, et al.
Pubblicazione: (2022)
di: Gruver, Nate, et al.
Pubblicazione: (2022)
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
di: Goldblum, Micah, et al.
Pubblicazione: (2023)
di: Goldblum, Micah, et al.
Pubblicazione: (2023)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
di: Kuang, Yilun, et al.
Pubblicazione: (2025)
di: Kuang, Yilun, et al.
Pubblicazione: (2025)
Large Language Models Are Zero-Shot Time Series Forecasters
di: Gruver, Nate, et al.
Pubblicazione: (2023)
di: Gruver, Nate, et al.
Pubblicazione: (2023)
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
di: Lotfi, Sanae, et al.
Pubblicazione: (2024)
di: Lotfi, Sanae, et al.
Pubblicazione: (2024)
Just How Flexible are Neural Networks in Practice?
di: Shwartz-Ziv, Ravid, et al.
Pubblicazione: (2024)
di: Shwartz-Ziv, Ravid, et al.
Pubblicazione: (2024)
Non-Vacuous Generalization Bounds for Large Language Models
di: Lotfi, Sanae, et al.
Pubblicazione: (2023)
di: Lotfi, Sanae, et al.
Pubblicazione: (2023)
Training Flexible Models of Genetic Variant Effects from Functional Annotations using Accelerated Linear Algebra
di: Amin, Alan N., et al.
Pubblicazione: (2025)
di: Amin, Alan N., et al.
Pubblicazione: (2025)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
di: Finzi, Marc, et al.
Pubblicazione: (2026)
di: Finzi, Marc, et al.
Pubblicazione: (2026)
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
Out-of-Distribution Detection Methods Answer the Wrong Questions
di: Li, Yucen Lily, et al.
Pubblicazione: (2025)
di: Li, Yucen Lily, et al.
Pubblicazione: (2025)
A Simple Baseline for Predicting Events with Auto-Regressive Tabular Transformers
di: Stein, Alex, et al.
Pubblicazione: (2024)
di: Stein, Alex, et al.
Pubblicazione: (2024)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
di: Marek, Martin, et al.
Pubblicazione: (2025)
di: Marek, Martin, et al.
Pubblicazione: (2025)
Compute-Optimal LLMs Provably Generalize Better With Scale
di: Finzi, Marc, et al.
Pubblicazione: (2025)
di: Finzi, Marc, et al.
Pubblicazione: (2025)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
di: Chen, Haoran, et al.
Pubblicazione: (2024)
di: Chen, Haoran, et al.
Pubblicazione: (2024)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
di: Qiu, Shikai, et al.
Pubblicazione: (2025)
AI versus AI in Financial Crimes and Detection: GenAI Crime Waves to Co-Evolutionary AI
di: Kurshan, Eren, et al.
Pubblicazione: (2024)
di: Kurshan, Eren, et al.
Pubblicazione: (2024)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
di: Jayawardhana, Mayuka, et al.
Pubblicazione: (2025)
di: Jayawardhana, Mayuka, et al.
Pubblicazione: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
di: Marek, Martin, et al.
Pubblicazione: (2026)
di: Marek, Martin, et al.
Pubblicazione: (2026)
BEDTime: A Unified Benchmark for Automatically Describing Time Series
di: Sen, Medhasweta, et al.
Pubblicazione: (2025)
di: Sen, Medhasweta, et al.
Pubblicazione: (2025)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
di: Zhang, Eva, et al.
Pubblicazione: (2024)
di: Zhang, Eva, et al.
Pubblicazione: (2024)
Dynamic Delayed Tree Expansion For Improved Multi-Path Speculative Decoding
di: Thomas, Rahul, et al.
Pubblicazione: (2026)
di: Thomas, Rahul, et al.
Pubblicazione: (2026)
Transferring Knowledge from Large Foundation Models to Small Downstream Models
di: Qiu, Shikai, et al.
Pubblicazione: (2024)
di: Qiu, Shikai, et al.
Pubblicazione: (2024)
L$^3$: Large Lookup Layers
di: Tseng, Albert, et al.
Pubblicazione: (2026)
di: Tseng, Albert, et al.
Pubblicazione: (2026)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
di: Souri, Hossein, et al.
Pubblicazione: (2024)
di: Souri, Hossein, et al.
Pubblicazione: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
di: Kapoor, Sanyam, et al.
Pubblicazione: (2024)
di: Kapoor, Sanyam, et al.
Pubblicazione: (2024)
Wohnungsnot, Geschlecht und Gesundheit
di: Finzi, Jan A.
Pubblicazione: (2023)
di: Finzi, Jan A.
Pubblicazione: (2023)
ESTRUCTURA DE PODER AL INTERIOR DE LA PAREJA Y DISCONFORT DE GÉNERO. REPRESENTACIONES DE LAS NORMAS DE GÉNERO EN LA FAMILIA CONTEMPORÁNEA ARGENTINA
di: Alejandra Martínez Finzi
Pubblicazione: (2012)
di: Alejandra Martínez Finzi
Pubblicazione: (2012)
Predicting the Performance of Black-box LLMs through Follow-up Queries
di: Sam, Dylan, et al.
Pubblicazione: (2025)
di: Sam, Dylan, et al.
Pubblicazione: (2025)
Diffusing Differentiable Representations
di: Savani, Yash, et al.
Pubblicazione: (2024)
di: Savani, Yash, et al.
Pubblicazione: (2024)
Privacy-Preserving Mechanisms Enable Cheap Verifiable Inference of LLMs
di: Pal, Arka, et al.
Pubblicazione: (2026)
di: Pal, Arka, et al.
Pubblicazione: (2026)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
di: Pal, Arka, et al.
Pubblicazione: (2025)
di: Pal, Arka, et al.
Pubblicazione: (2025)
Preserving Cross-Modal Consistency for CLIP-based Class-Incremental Learning
di: Chen, Haoran, et al.
Pubblicazione: (2025)
di: Chen, Haoran, et al.
Pubblicazione: (2025)
Pointer: Linear-Complexity Long-Range Modeling without Pre-training
di: Li, Zixi
Pubblicazione: (2025)
di: Li, Zixi
Pubblicazione: (2025)
Zero-shot Multivariate Time Series Forecasting Using Tabular Prior Fitted Networks
di: Jayawardhana, Mayuka, et al.
Pubblicazione: (2026)
di: Jayawardhana, Mayuka, et al.
Pubblicazione: (2026)
TimeSqueeze: Dynamic Patching for Efficient Time Series Forecasting
di: Ankireddy, Sravan Kumar, et al.
Pubblicazione: (2026)
di: Ankireddy, Sravan Kumar, et al.
Pubblicazione: (2026)
Influencia de la presión y la temperatura en la reacción de hidroformilación en medio bifásico del 1-hexeno en régimen continuo con el complejo catalítico [(µ-Pz)(CO)(TFFTS)Rh]2 (I)
di: Arnoldo Bruss
Pubblicazione: (2006)
di: Arnoldo Bruss
Pubblicazione: (2006)
Modulation by context of a scene in monkey anterior inferotemporal cortex during a saccadic eye movement task
di: Bruss Lima
Pubblicazione: (2003)
di: Bruss Lima
Pubblicazione: (2003)
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Compute Better Spent: Replacing Dense Layers with Structured Matrices
di: Qiu, Shikai, et al.
Pubblicazione: (2024) -
The Lie Derivative for Measuring Learned Equivariance
di: Gruver, Nate, et al.
Pubblicazione: (2022) -
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
di: Goldblum, Micah, et al.
Pubblicazione: (2023) -
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
di: Kuang, Yilun, et al.
Pubblicazione: (2025) -
Large Language Models Are Zero-Shot Time Series Forecasters
di: Gruver, Nate, et al.
Pubblicazione: (2023)