Gespeichert in:
| Hauptverfasser: | Shamshoum, Yara, Hodos, Nitzan, Sieradzki, Yuval, Schuster, Assaf |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.15352 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DNCs Require More Planning Steps
von: Shamshoum, Yara, et al.
Veröffentlicht: (2024)
von: Shamshoum, Yara, et al.
Veröffentlicht: (2024)
QKV Projections Require a Fraction of Their Memory
von: Khalaf, Malik, et al.
Veröffentlicht: (2025)
von: Khalaf, Malik, et al.
Veröffentlicht: (2025)
CompAct: Compressing Retrieved Documents Actively for Question Answering
von: Yoon, Chanwoong, et al.
Veröffentlicht: (2024)
von: Yoon, Chanwoong, et al.
Veröffentlicht: (2024)
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
von: Solgi, Ryan, et al.
Veröffentlicht: (2025)
von: Solgi, Ryan, et al.
Veröffentlicht: (2025)
Memory-Efficient LLM Training with Online Subspace Descent
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
ActTail: Global Activation Sparsity in Large Language Models
von: Hou, Wenwen, et al.
Veröffentlicht: (2026)
von: Hou, Wenwen, et al.
Veröffentlicht: (2026)
SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM Training
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
von: Refael, Yehonathan, et al.
Veröffentlicht: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
von: Ping, Bowen, et al.
Veröffentlicht: (2026)
von: Ping, Bowen, et al.
Veröffentlicht: (2026)
OneComp: One-Line Revolution for Generative AI Model Compression
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2026)
von: Ichikawa, Yuma, et al.
Veröffentlicht: (2026)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
von: Ackerman, Samuel, et al.
Veröffentlicht: (2025)
von: Ackerman, Samuel, et al.
Veröffentlicht: (2025)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
Prompt Curriculum Learning for Efficient LLM Post-Training
von: Gao, Zhaolin, et al.
Veröffentlicht: (2025)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2025)
Learning is Forgetting: LLM Training As Lossy Compression
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
von: Conklin, Henry C., et al.
Veröffentlicht: (2026)
LoMA: Lossless Compressed Memory Attention
von: Wang, Yumeng, et al.
Veröffentlicht: (2024)
von: Wang, Yumeng, et al.
Veröffentlicht: (2024)
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
Think Before You Act: Decision Transformers with Working Memory
von: Kang, Jikun, et al.
Veröffentlicht: (2023)
von: Kang, Jikun, et al.
Veröffentlicht: (2023)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
Compressed Context Memory For Online Language Model Interaction
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2023)
von: Kim, Jang-Hyun, et al.
Veröffentlicht: (2023)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark
von: Zhang, Yihua, et al.
Veröffentlicht: (2024)
von: Zhang, Yihua, et al.
Veröffentlicht: (2024)
LCSB: Layer-Cyclic Selective Backpropagation for Memory-Efficient On-Device LLM Fine-Tuning
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
AdaFRUGAL: Adaptive Memory-Efficient Training with Dynamic Control
von: Bui, Quang-Hung, et al.
Veröffentlicht: (2025)
von: Bui, Quang-Hung, et al.
Veröffentlicht: (2025)
Training LLMs over Neurally Compressed Text
von: Lester, Brian, et al.
Veröffentlicht: (2024)
von: Lester, Brian, et al.
Veröffentlicht: (2024)
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
Adaptive Querying with AI Persona Priors
von: Wang, Kaizheng, et al.
Veröffentlicht: (2026)
von: Wang, Kaizheng, et al.
Veröffentlicht: (2026)
Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
First Activations Matter: Training-Free Methods for Dynamic Activation in Large Language Models
von: Ma, Chi, et al.
Veröffentlicht: (2024)
von: Ma, Chi, et al.
Veröffentlicht: (2024)
Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression
von: Trukhina, Natalia, et al.
Veröffentlicht: (2026)
von: Trukhina, Natalia, et al.
Veröffentlicht: (2026)
Evaluating Memory Structure in LLM Agents
von: Shutova, Alina, et al.
Veröffentlicht: (2026)
von: Shutova, Alina, et al.
Veröffentlicht: (2026)
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
von: Gera, Ariel, et al.
Veröffentlicht: (2026)
von: Gera, Ariel, et al.
Veröffentlicht: (2026)
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
Scaling with Collapse: Efficient and Predictable Training of LLM Families
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
von: Bergsma, Shane, et al.
Veröffentlicht: (2025)
LLM Router: Rethinking Routing with Prefill Activations
von: Varshney, Tanay, et al.
Veröffentlicht: (2026)
von: Varshney, Tanay, et al.
Veröffentlicht: (2026)
Prior-Informed Zeroth-Order Optimization with Adaptive Direction Alignment for Memory-Efficient LLM Fine-Tuning
von: Jin, Feihu, et al.
Veröffentlicht: (2026)
von: Jin, Feihu, et al.
Veröffentlicht: (2026)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient Training
von: Li, Yanyi, et al.
Veröffentlicht: (2026)
von: Li, Yanyi, et al.
Veröffentlicht: (2026)
DataComp-LM: In search of the next generation of training sets for language models
von: Li, Jeffrey, et al.
Veröffentlicht: (2024)
von: Li, Jeffrey, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DNCs Require More Planning Steps
von: Shamshoum, Yara, et al.
Veröffentlicht: (2024) -
QKV Projections Require a Fraction of Their Memory
von: Khalaf, Malik, et al.
Veröffentlicht: (2025) -
CompAct: Compressing Retrieved Documents Actively for Question Answering
von: Yoon, Chanwoong, et al.
Veröffentlicht: (2024) -
Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
von: Solgi, Ryan, et al.
Veröffentlicht: (2025) -
Memory-Efficient LLM Training with Online Subspace Descent
von: Liang, Kaizhao, et al.
Veröffentlicht: (2024)