Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kinas, Remigiusz, Kiszczak, Paweł, Perez, Sergio P., Ociepa, Krzysztof, Flis, Łukasz, Wróbel, Krzysztof, Gwoździej, Adrian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2024)
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2024)
Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2026)
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2026)
Bielik 11B v3: Multilingual Large Language Model for European Languages
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
Bielik 11B v2 Technical Report
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
Bielik v3 Small: Technical Report
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
von: Wróbel, Krzysztof, et al.
Veröffentlicht: (2026)
von: Wróbel, Krzysztof, et al.
Veröffentlicht: (2026)
PL-Guard: Benchmarking Language Model Safety for Polish
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
Large Language Models for Biomedical Article Classification
von: Proboszcz, Jakub, et al.
Veröffentlicht: (2026)
von: Proboszcz, Jakub, et al.
Veröffentlicht: (2026)
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
von: Christop, Iwona, et al.
Veröffentlicht: (2026)
von: Christop, Iwona, et al.
Veröffentlicht: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
Cross-Lingual Generalization and Compression: From Language-Specific to Shared Neurons
von: Riemenschneider, Frederick, et al.
Veröffentlicht: (2025)
von: Riemenschneider, Frederick, et al.
Veröffentlicht: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
von: Chang, Hoyeon, et al.
Veröffentlicht: (2024)
von: Chang, Hoyeon, et al.
Veröffentlicht: (2024)
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2023)
von: Park, Seungcheol, et al.
Veröffentlicht: (2023)
The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
von: Sauter, Adrian, et al.
Veröffentlicht: (2025)
von: Sauter, Adrian, et al.
Veröffentlicht: (2025)
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
von: Labadie-Tamayo, Roberto, et al.
Veröffentlicht: (2025)
SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models
von: Rubinstein, Beny, et al.
Veröffentlicht: (2026)
von: Rubinstein, Beny, et al.
Veröffentlicht: (2026)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
Right for Right Reasons: Large Language Models for Verifiable Commonsense Knowledge Graph Question Answering
von: Toroghi, Armin, et al.
Veröffentlicht: (2024)
von: Toroghi, Armin, et al.
Veröffentlicht: (2024)
Streamlining Redundant Layers to Compress Large Language Models
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
von: Chen, Xiaodong, et al.
Veröffentlicht: (2024)
Truth as a Compression Artifact in Language Model Training
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
Distilling Large Language Models for Efficient Clinical Information Extraction
von: Vedula, Karthik S., et al.
Veröffentlicht: (2024)
von: Vedula, Karthik S., et al.
Veröffentlicht: (2024)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
Does Localization Inform Unlearning? A Rigorous Examination of Local Parameter Attribution for Knowledge Unlearning in Language Models
von: Lee, Hwiyeong, et al.
Veröffentlicht: (2025)
von: Lee, Hwiyeong, et al.
Veröffentlicht: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
Memory Bank Compression for Continual Adaptation of Large Language Models
von: Katraouras, Thomas, et al.
Veröffentlicht: (2026)
von: Katraouras, Thomas, et al.
Veröffentlicht: (2026)
Distilling Robustness into Natural Language Inference Models with Domain-Targeted Augmentation
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
von: Stacey, Joe, et al.
Veröffentlicht: (2023)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Engineering A Large Language Model From Scratch
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
$\rm SP^3$: Enhancing Structured Pruning via PCA Projection
von: Hu, Yuxuan, et al.
Veröffentlicht: (2023)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2023)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning
von: Zou, Run, et al.
Veröffentlicht: (2026)
von: Zou, Run, et al.
Veröffentlicht: (2026)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Semantic Query Processing in a Scholarly Knowledge Graph
von: Jia, Runsong, et al.
Veröffentlicht: (2024)
von: Jia, Runsong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2024) -
Advancing Polish Language Modeling through Tokenizer Optimization in the Bielik v3 7B and 11B Series
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2026) -
Bielik 11B v3: Multilingual Large Language Model for European Languages
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025) -
Bielik 11B v2 Technical Report
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025) -
Bielik v3 Small: Technical Report
von: Ociepa, Krzysztof, et al.
Veröffentlicht: (2025)