Salvato in:
| Autori principali: | Hay, Tamir David, Wolf, Lior |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2401.12819 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Error Correction Code Transformers
di: Levy, Matan, et al.
Pubblicazione: (2024)
di: Levy, Matan, et al.
Pubblicazione: (2024)
Factor Graph Optimization of Error-Correcting Codes for Belief Propagation Decoding
di: Choukroun, Yoni, et al.
Pubblicazione: (2024)
di: Choukroun, Yoni, et al.
Pubblicazione: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
di: Katz, Shahar, et al.
Pubblicazione: (2024)
di: Katz, Shahar, et al.
Pubblicazione: (2024)
Module-Aware Parameter-Efficient Machine Unlearning on Transformers
di: Bao, Wenjie, et al.
Pubblicazione: (2025)
di: Bao, Wenjie, et al.
Pubblicazione: (2025)
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
di: Indelman, Hedda Cohen, et al.
Pubblicazione: (2024)
di: Indelman, Hedda Cohen, et al.
Pubblicazione: (2024)
Evaluating Neural Networks for Early Maritime Threat Detection
di: Tella, Dhanush, et al.
Pubblicazione: (2024)
di: Tella, Dhanush, et al.
Pubblicazione: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
di: Kamigaito, Hidetaka, et al.
Pubblicazione: (2025)
di: Kamigaito, Hidetaka, et al.
Pubblicazione: (2025)
TWISTED-RL: Hierarchical Skilled Agents for Knot-Tying without Human Demonstrations
di: Freund, Guy, et al.
Pubblicazione: (2026)
di: Freund, Guy, et al.
Pubblicazione: (2026)
Finding Clustering Algorithms in the Transformer Architecture
di: Clarkson, Kenneth L., et al.
Pubblicazione: (2025)
di: Clarkson, Kenneth L., et al.
Pubblicazione: (2025)
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
di: Hu, Jerry Yao-Chieh, et al.
Pubblicazione: (2024)
Understanding and Guiding Layer Placement in Parameter-Efficient Fine-Tuning of Large Language Models
di: Xu, Yichen, et al.
Pubblicazione: (2026)
di: Xu, Yichen, et al.
Pubblicazione: (2026)
Set-based Neural Network Encoding Without Weight Tying
di: Andreis, Bruno, et al.
Pubblicazione: (2023)
di: Andreis, Bruno, et al.
Pubblicazione: (2023)
Anomaly Detection with Variance Stabilized Density Estimation
di: Rozner, Amit, et al.
Pubblicazione: (2023)
di: Rozner, Amit, et al.
Pubblicazione: (2023)
Deep Active Speech Cancellation with Mamba-Masking Network
di: Mishaly, Yehuda, et al.
Pubblicazione: (2025)
di: Mishaly, Yehuda, et al.
Pubblicazione: (2025)
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
di: Eisenstadt, Roy, et al.
Pubblicazione: (2025)
di: Eisenstadt, Roy, et al.
Pubblicazione: (2025)
Echo-LoRA: Parameter-Efficient Fine-Tuning via Cross-Layer Representation Injection
di: Peng, Yihang, et al.
Pubblicazione: (2026)
di: Peng, Yihang, et al.
Pubblicazione: (2026)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
di: Cohen, Lior, et al.
Pubblicazione: (2025)
di: Cohen, Lior, et al.
Pubblicazione: (2025)
Partial Parameter Updates for Efficient Distributed Training
di: Filippova, Anastasiia, et al.
Pubblicazione: (2025)
di: Filippova, Anastasiia, et al.
Pubblicazione: (2025)
Parameter-Efficient Fine-Tuning with Discrete Fourier Transform
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
Is there Value in Reinforcement Learning?
di: Fox, Lior, et al.
Pubblicazione: (2025)
di: Fox, Lior, et al.
Pubblicazione: (2025)
Communication Dynamics Neural Networks: FFT-Diagonalized Layers for Improved Hessian Conditioning at Reduced Parameter Count
di: Pan, Lurong
Pubblicazione: (2026)
di: Pan, Lurong
Pubblicazione: (2026)
Non-Uniform Parameter-Wise Model Merging
di: Camacho, Albert Manuel Orozco, et al.
Pubblicazione: (2024)
di: Camacho, Albert Manuel Orozco, et al.
Pubblicazione: (2024)
PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers in a resource-limited Context
di: Augustin, Maximilian, et al.
Pubblicazione: (2024)
di: Augustin, Maximilian, et al.
Pubblicazione: (2024)
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
di: Ham, Seokil, et al.
Pubblicazione: (2024)
di: Ham, Seokil, et al.
Pubblicazione: (2024)
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
di: Zafar, Oz, et al.
Pubblicazione: (2024)
di: Zafar, Oz, et al.
Pubblicazione: (2024)
Parameter-Efficient Fine-Tuning for HAR: Integrating LoRA and QLoRA into Transformer Models
di: Seregina, Irina, et al.
Pubblicazione: (2025)
di: Seregina, Irina, et al.
Pubblicazione: (2025)
DeciMamba: Exploring the Length Extrapolation Potential of Mamba
di: Ben-Kish, Assaf, et al.
Pubblicazione: (2024)
di: Ben-Kish, Assaf, et al.
Pubblicazione: (2024)
Overflow Prevention Enhances Long-Context Recurrent LLMs
di: Ben-Kish, Assaf, et al.
Pubblicazione: (2025)
di: Ben-Kish, Assaf, et al.
Pubblicazione: (2025)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
di: Wu, Qitian, et al.
Pubblicazione: (2024)
di: Wu, Qitian, et al.
Pubblicazione: (2024)
Hybrid Dynamic Pruning: A Pathway to Efficient Transformer Inference
di: Jaradat, Ghadeer, et al.
Pubblicazione: (2024)
di: Jaradat, Ghadeer, et al.
Pubblicazione: (2024)
Solo Connection: A Parameter Efficient Fine-Tuning Technique for Transformers
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
Enhancing GNNs with Architecture-Agnostic Graph Transformations: A Systematic Analysis
di: Li, Zhifei, et al.
Pubblicazione: (2024)
di: Li, Zhifei, et al.
Pubblicazione: (2024)
Identifying Bias in Deep Neural Networks Using Image Transforms
di: Erukude, Sai Teja, et al.
Pubblicazione: (2024)
di: Erukude, Sai Teja, et al.
Pubblicazione: (2024)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
di: Kapadia, Shashank, et al.
Pubblicazione: (2026)
di: Kapadia, Shashank, et al.
Pubblicazione: (2026)
Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models
di: Guo, Yongxin, et al.
Pubblicazione: (2024)
di: Guo, Yongxin, et al.
Pubblicazione: (2024)
PAINET: A Principled Efficient Transformer for 3D Dynamics Modeling
di: Yang, Kai, et al.
Pubblicazione: (2025)
di: Yang, Kai, et al.
Pubblicazione: (2025)
Unified Parameter-Efficient Unlearning for LLMs
di: Ding, Chenlu, et al.
Pubblicazione: (2024)
di: Ding, Chenlu, et al.
Pubblicazione: (2024)
Transformer Normalisation Layers and the Independence of Semantic Subspaces
di: Menary, Stephen, et al.
Pubblicazione: (2024)
di: Menary, Stephen, et al.
Pubblicazione: (2024)
Layer Specialization Underlying Compositional Reasoning in Transformers
di: Liu, Jing
Pubblicazione: (2025)
di: Liu, Jing
Pubblicazione: (2025)
DLO: Dynamic Layer Operation for Efficient Vertical Scaling of LLMs
di: Tan, Zhen, et al.
Pubblicazione: (2024)
di: Tan, Zhen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Accelerating Error Correction Code Transformers
di: Levy, Matan, et al.
Pubblicazione: (2024) -
Factor Graph Optimization of Error-Correcting Codes for Belief Propagation Decoding
di: Choukroun, Yoni, et al.
Pubblicazione: (2024) -
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
di: Katz, Shahar, et al.
Pubblicazione: (2024) -
Module-Aware Parameter-Efficient Machine Unlearning on Transformers
di: Bao, Wenjie, et al.
Pubblicazione: (2025) -
On The Statistical Representation Properties Of The Perturb-Softmax And The Perturb-Argmax Probability Distributions
di: Indelman, Hedda Cohen, et al.
Pubblicazione: (2024)