Strategic Fusion Optimizes Transformer Compression
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Rahman, Md Shoaibur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025)
Emotion Detection From Social Media Posts
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023)
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
von: Zhang, Jing, et al.
Veröffentlicht: (2024)
von: Zhang, Jing, et al.
Veröffentlicht: (2024)
PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models
von: Yu, Ye, et al.
Veröffentlicht: (2025)
von: Yu, Ye, et al.
Veröffentlicht: (2025)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
von: Rahman, Subhey Sadi, et al.
Veröffentlicht: (2025)
Implicit Strategic Optimization: Rethinking Long-Horizon Decision-Making in Adversarial Poker Environments
von: Xia, Boyang, et al.
Veröffentlicht: (2026)
von: Xia, Boyang, et al.
Veröffentlicht: (2026)
MultiSoc-4D: A Benchmark for Diagnosing Instruction-Induced Label Collapse in Closed-Set LLM Annotation of Bengali Social Media
von: Pramanik, Souvik, et al.
Veröffentlicht: (2026)
von: Pramanik, Souvik, et al.
Veröffentlicht: (2026)
Momentum Streams for Optimizer-Inspired Transformers
von: Gai, Jingchu, et al.
Veröffentlicht: (2026)
von: Gai, Jingchu, et al.
Veröffentlicht: (2026)
What Differentiates Educational Literature? A Multimodal Fusion Approach of Transformers and Computational Linguistics
von: Bird, Jordan J.
Veröffentlicht: (2024)
von: Bird, Jordan J.
Veröffentlicht: (2024)
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
von: Girija, Sanjay Surendranath, et al.
Veröffentlicht: (2025)
von: Girija, Sanjay Surendranath, et al.
Veröffentlicht: (2025)
LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
von: Li, Guangyan, et al.
Veröffentlicht: (2024)
von: Li, Guangyan, et al.
Veröffentlicht: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models
von: Islam, Shayekh Bin, et al.
Veröffentlicht: (2024)
von: Islam, Shayekh Bin, et al.
Veröffentlicht: (2024)
ROCKET: Rapid Optimization via Calibration-guided Knapsack Enhanced Truncation for Efficient Model Compression
von: Ali, Ammar, et al.
Veröffentlicht: (2026)
von: Ali, Ammar, et al.
Veröffentlicht: (2026)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
von: Wang, Xinyuan, et al.
Veröffentlicht: (2026)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2026)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
von: Chen, Yifang, et al.
Veröffentlicht: (2025)
Beyond Fertility: Analyzing STRR as a Metric for Multilingual Tokenization Evaluation
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2025)
von: Nayeem, Mir Tafseer, et al.
Veröffentlicht: (2025)
DACP: Domain-Adaptive Continual Pre-Training of Large Language Models for Phone Conversation Summarization
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
von: Fu, Xue-Yong, et al.
Veröffentlicht: (2025)
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2026)
Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning
von: Rahman, Md Mushfiqur, et al.
Veröffentlicht: (2024)
von: Rahman, Md Mushfiqur, et al.
Veröffentlicht: (2024)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
von: Choi, Euntae, et al.
Veröffentlicht: (2025)
H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
von: Tekin, Selim Furkan, et al.
Veröffentlicht: (2024)
AI Knowledge Assist: An Automated Approach for the Creation of Knowledge Bases for Conversational AI Agents
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2024)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
von: Qiao, Aurick, et al.
Veröffentlicht: (2024)
FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model Fusion
von: Tang, Anke, et al.
Veröffentlicht: (2024)
von: Tang, Anke, et al.
Veröffentlicht: (2024)
PrismRAG: Boosting RAG Factuality with Distractor Resilience and Strategized Reasoning
von: Kachuee, Mohammad, et al.
Veröffentlicht: (2025)
von: Kachuee, Mohammad, et al.
Veröffentlicht: (2025)
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
von: Zuo, Bowen, et al.
Veröffentlicht: (2025)
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
von: Yen, Jui-Nan, et al.
Veröffentlicht: (2024)
von: Yen, Jui-Nan, et al.
Veröffentlicht: (2024)
Linear Chain Transformation: Expanding Optimization Dynamics for Fine-Tuning Large Language Models
von: Wang, Yulong, et al.
Veröffentlicht: (2024)
von: Wang, Yulong, et al.
Veröffentlicht: (2024)
On the Compressibility of Quantized Large Language Models
von: Mao, Yu, et al.
Veröffentlicht: (2024)
von: Mao, Yu, et al.
Veröffentlicht: (2024)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary
von: Zheng, Qirui, et al.
Veröffentlicht: (2025)
von: Zheng, Qirui, et al.
Veröffentlicht: (2025)
A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
von: Scheurer, Jérémy, et al.
Veröffentlicht: (2023)
Fast Vocabulary Transfer for Language Model Compression
von: Gee, Leonidas, et al.
Veröffentlicht: (2024)
von: Gee, Leonidas, et al.
Veröffentlicht: (2024)
Communication Compression for Tensor Parallel LLM Inference
von: Hansen-Palmus, Jan, et al.
Veröffentlicht: (2024)
von: Hansen-Palmus, Jan, et al.
Veröffentlicht: (2024)
Learning to Compress Prompt in Natural Language Formats
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Projected Compression: Trainable Projection for Efficient Transformer Compression
von: Stefaniak, Maciej, et al.
Veröffentlicht: (2025) -
Emotion Detection From Social Media Posts
von: Rahman, Md Mahbubur, et al.
Veröffentlicht: (2023) -
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization
von: Rahman, Ben
Veröffentlicht: (2025) -
SEE: Sememe Entanglement Encoding for Transformer-bases Models Compression
von: Zhang, Jing, et al.
Veröffentlicht: (2024) -
PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models
von: Yu, Ye, et al.
Veröffentlicht: (2025)