OpenBA-V2: Reaching 77.3% High Compression Ratio with Fast Multi-Stage Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Qiao, Dan, Su, Yi, Wang, Pinzheng, Ye, Jing, Xie, Wenjing, Zhou, Yuechi, Ding, Yuyang, Tang, Zecheng, Wang, Jikai, Ji, Yixin, Wang, Yue, Guo, Pei, Sun, Zechen, Zhang, Zikang, Li, Juntao, Chao, Pingfu, Chen, Wenliang, Fu, Guohong, Zhou, Guodong, Zhu, Qiaoming, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenBA: An Open-sourced 15B Bilingual Asymmetric seq2seq Model Pre-trained from Scratch
by: Li, Juntao, et al.
Published: (2023)
by: Li, Juntao, et al.
Published: (2023)
Revealing and Mitigating Over-Attention in Knowledge Editing
by: Wang, Pinzheng, et al.
Published: (2025)
by: Wang, Pinzheng, et al.
Published: (2025)
LOGO -- Long cOntext aliGnment via efficient preference Optimization
by: Tang, Zecheng, et al.
Published: (2024)
by: Tang, Zecheng, et al.
Published: (2024)
Rethinking Negative Instances for Generative Named Entity Recognition
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
CMD: a framework for Context-aware Model self-Detoxification
by: Tang, Zecheng, et al.
Published: (2023)
by: Tang, Zecheng, et al.
Published: (2023)
Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
by: Wang, Pinzheng, et al.
Published: (2025)
by: Wang, Pinzheng, et al.
Published: (2025)
Towards DS-NER: Unveiling and Addressing Latent Noise in Distant Annotations
by: Ding, Yuyang, et al.
Published: (2025)
by: Ding, Yuyang, et al.
Published: (2025)
Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
by: Xiang, Yang, et al.
Published: (2025)
by: Xiang, Yang, et al.
Published: (2025)
LongFlow: Efficient KV Cache Compression for Reasoning Models
by: Su, Yi, et al.
Published: (2026)
by: Su, Yi, et al.
Published: (2026)
Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
$A^3$: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving
by: Zhou, Yuechi, et al.
Published: (2025)
by: Zhou, Yuechi, et al.
Published: (2025)
$\textbf{Re}^{2}$: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving
by: Wang, Pinzheng, et al.
Published: (2026)
by: Wang, Pinzheng, et al.
Published: (2026)
Beware of Calibration Data for Pruning Large Language Models
by: Ji, Yixin, et al.
Published: (2024)
by: Ji, Yixin, et al.
Published: (2024)
Resilient Cooperative NE Control for High‐Order Nonlinear MASs Under FDI and DoS Attacks
by: Xia Zhou, et al.
Published: (2026)
by: Xia Zhou, et al.
Published: (2026)
Accurate KV Cache Quantization with Outlier Tokens Tracing
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
CaliDrop: KV Cache Compression with Calibration
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
773 | EFFICACY AND SAFETY OF SELINEXOR, CARFILZOMIB, POMALIDOMIDE AND DEXAMETHASONE IN THE TREATMENT OF MM WITH EXTRAMEDULLARY DISEASE: A PROSPECTIVE MULTI‐CENTER STUDY
by: H. Zhou, et al.
Published: (2025)
by: H. Zhou, et al.
Published: (2025)
RuPLaR : Efficient Latent Compression of LLM Reasoning Chains with Rule-Based Priors From Multi-Step to One-Step
by: Luo, Xiaocheng, et al.
Published: (2026)
by: Luo, Xiaocheng, et al.
Published: (2026)
Optimization Model of Steel‐Prestressed Concrete Hybrid Wind Turbine Tower: Using a Combined Differential Whale Optimization Algorithm
by: Wei Xu, et al.
Published: (2025)
by: Wei Xu, et al.
Published: (2025)
L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding?
by: Tang, Zecheng, et al.
Published: (2024)
by: Tang, Zecheng, et al.
Published: (2024)
Efficient Reasoning for LLMs through Speculative Chain-of-Thought
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers
by: Wang, Haochuan Kevin, et al.
Published: (2026)
by: Wang, Haochuan Kevin, et al.
Published: (2026)
Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers
by: Tang, Zecheng, et al.
Published: (2026)
by: Tang, Zecheng, et al.
Published: (2026)
Seabed photographs taken along OFOS profile MSM77_3-5 during RV MARIA S. MERIAN cruise MSM77
by: Boehringer, Lilian, et al.
Published: (2026)
by: Boehringer, Lilian, et al.
Published: (2026)
LOOM-Scope: a comprehensive and efficient LOng-cOntext Model evaluation framework
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
EviRerank: Adaptive Evidence Construction for Long-Document LLM Reranking
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
The minimal model of Rota-Baxter operad with arbitrary weight
by: Wang, Kai, et al.
Published: (2022)
by: Wang, Kai, et al.
Published: (2022)
Neural Scaling Laws of Deep ReLU and Deep Operator Network: A Theoretical Study
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Coefficient-to-Basis Network: A Fine-Tunable Operator Learning Framework for Inverse Problems with Adaptive Discretizations and Theoretical Guarantees
by: Zhang, Zecheng, et al.
Published: (2025)
by: Zhang, Zecheng, et al.
Published: (2025)
Effect of protection zone on the dynamics of a diffusion-advection population-toxicant model
by: Gao, Jing, et al.
Published: (2025)
by: Gao, Jing, et al.
Published: (2025)
Grazing duration and intensity modulate vegetation dynamics in semi-arid ecosystems with seasonal succession
by: Gan, Junhong, et al.
Published: (2025)
by: Gan, Junhong, et al.
Published: (2025)
Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
by: Rong, Yiming, et al.
Published: (2025)
by: Rong, Yiming, et al.
Published: (2025)
Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure
by: Wang, Jikai, et al.
Published: (2024)
by: Wang, Jikai, et al.
Published: (2024)
Image anomaly detection and prediction scheme based on SSA optimized ResNet50-BiGRU model
by: Wan, Qianhui, et al.
Published: (2024)
by: Wan, Qianhui, et al.
Published: (2024)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
by: Chen, Peiyuan, et al.
Published: (2024)
by: Chen, Peiyuan, et al.
Published: (2024)
Theoretical Analysis of Meta Reinforcement Learning: Generalization Bounds and Convergence Guarantees
by: Wang, Cangqing, et al.
Published: (2024)
by: Wang, Cangqing, et al.
Published: (2024)
Deep Analysis of Time Series Data for Smart Grid Startup Strategies: A Transformer-LSTM-PSO Model Approach
by: Zhang, Zecheng
Published: (2024)
by: Zhang, Zecheng
Published: (2024)
MODNO: Multi Operator Learning With Distributed Neural Operators
by: Zhang, Zecheng
Published: (2024)
by: Zhang, Zecheng
Published: (2024)
MemLong: Memory-Augmented Retrieval for Long Text Modeling
by: Liu, Weijie, et al.
Published: (2024)
by: Liu, Weijie, et al.
Published: (2024)
Similar Items
-
OpenBA: An Open-sourced 15B Bilingual Asymmetric seq2seq Model Pre-trained from Scratch
by: Li, Juntao, et al.
Published: (2023) -
Revealing and Mitigating Over-Attention in Knowledge Editing
by: Wang, Pinzheng, et al.
Published: (2025) -
LOGO -- Long cOntext aliGnment via efficient preference Optimization
by: Tang, Zecheng, et al.
Published: (2024) -
Rethinking Negative Instances for Generative Named Entity Recognition
by: Ding, Yuyang, et al.
Published: (2024) -
CMD: a framework for Context-aware Model self-Detoxification
by: Tang, Zecheng, et al.
Published: (2023)