Krony-PT: GPT2 compressed with Kronecker Products
Fuente:
arXiv
Saved in:
| Main Authors: | Ayad, Mohamed Ayoub Ben, Mitrovic, Jelena, Granitzer, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Challenges and Considerations in Annotating Legal Data: A Comprehensive Overview
by: Darji, Harshil, et al.
Published: (2024)
by: Darji, Harshil, et al.
Published: (2024)
Compressed Concatenation of Small Embedding Models
by: Ayad, Mohamed Ayoub Ben, et al.
Published: (2025)
by: Ayad, Mohamed Ayoub Ben, et al.
Published: (2025)
Technical Report: Impact of Position Bias on Language Models in Token Classification
by: Amor, Mehdi Ben, et al.
Published: (2023)
by: Amor, Mehdi Ben, et al.
Published: (2023)
Computational Approaches to the Detection of Lesser-Known Rhetorical Figures: A Systematic Survey and Research Challenges
by: Kühn, Ramona, et al.
Published: (2024)
by: Kühn, Ramona, et al.
Published: (2024)
Enhancing Rhetorical Figure Annotation: An Ontology-Based Web Application with RAG Integration
by: Kühn, Ramona, et al.
Published: (2024)
by: Kühn, Ramona, et al.
Published: (2024)
KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices
by: Zhou, Wuyang, et al.
Published: (2026)
by: Zhou, Wuyang, et al.
Published: (2026)
SUDS: A Strategy for Unsupervised Drift Sampling
by: Fellicious, Christofer, et al.
Published: (2024)
by: Fellicious, Christofer, et al.
Published: (2024)
WebFAQ 2.0: A Multilingual QA Dataset with Mined Hard Negatives for Dense Retrieval
by: Dinzinger, Michael, et al.
Published: (2026)
by: Dinzinger, Michael, et al.
Published: (2026)
WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval
by: Dinzinger, Michael, et al.
Published: (2025)
by: Dinzinger, Michael, et al.
Published: (2025)
Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
by: Kurochkin, Vadim, et al.
Published: (2025)
by: Kurochkin, Vadim, et al.
Published: (2025)
MoKA: Mixture of Kronecker Adapters
by: Sadeghi, Mohammadreza, et al.
Published: (2025)
by: Sadeghi, Mohammadreza, et al.
Published: (2025)
Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations
by: Dang, Dang H., et al.
Published: (2026)
by: Dang, Dang H., et al.
Published: (2026)
TransactionGPT
by: Dou, Yingtong, et al.
Published: (2025)
by: Dou, Yingtong, et al.
Published: (2025)
Benevolent Dictators? On LLM Agent Behavior in Dictator Games
by: Einwiller, Andreas, et al.
Published: (2025)
by: Einwiller, Andreas, et al.
Published: (2025)
MiniGPT: Rebuilding GPT from First Principles
by: Joseph, Jibin
Published: (2026)
by: Joseph, Jibin
Published: (2026)
ZzzGPT: An Interactive GPT Approach to Enhance Sleep Quality
by: Khaokaew, Yonchanok, et al.
Published: (2023)
by: Khaokaew, Yonchanok, et al.
Published: (2023)
GPT Meets Graphs and KAN Splines: Testing Novel Frameworks on Multitask Fine-Tuned GPT-2 with LoRA
by: Bo, Gabriel, et al.
Published: (2025)
by: Bo, Gabriel, et al.
Published: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
by: Africa, David Demitri
Published: (2025)
by: Africa, David Demitri
Published: (2025)
ALIEN: Aligned Entropy Head for Improving Uncertainty Estimation of LLMs
by: Zabolotnyi, Artem, et al.
Published: (2025)
by: Zabolotnyi, Artem, et al.
Published: (2025)
M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
by: Wang, Taowen, et al.
Published: (2024)
by: Wang, Taowen, et al.
Published: (2024)
You can remove GPT2's LayerNorm by fine-tuning
by: Heimersheim, Stefan
Published: (2024)
by: Heimersheim, Stefan
Published: (2024)
OwlerLite: Scope- and Freshness-Aware Web Retrieval for LLM Assistants
by: Zerhoudi, Saber, et al.
Published: (2026)
by: Zerhoudi, Saber, et al.
Published: (2026)
Single layer tiny Co$^4$ outpaces GPT-2 and GPT-BERT
by: Zain, Noor Ul, et al.
Published: (2025)
by: Zain, Noor Ul, et al.
Published: (2025)
On Training Data Influence of GPT Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
ADaPT: As-Needed Decomposition and Planning with Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
ScaffoldGPT: A Scaffold-based GPT Model for Drug Optimization
by: Liu, Xuefeng, et al.
Published: (2025)
by: Liu, Xuefeng, et al.
Published: (2025)
Using ChatGPT for Data Science Analyses
by: Evkaya, Ozan, et al.
Published: (2024)
by: Evkaya, Ozan, et al.
Published: (2024)
HiGPT: Heterogeneous Graph Language Model
by: Tang, Jiabin, et al.
Published: (2024)
by: Tang, Jiabin, et al.
Published: (2024)
In-Context Learning and Fine-Tuning GPT for Argument Mining
by: Cabessa, Jérémie, et al.
Published: (2024)
by: Cabessa, Jérémie, et al.
Published: (2024)
Cross-Language Assessment of Mathematical Capability of ChatGPT
by: Sathe, Gargi, et al.
Published: (2024)
by: Sathe, Gargi, et al.
Published: (2024)
On Sarcasm Detection with OpenAI GPT-based Models
by: Gole, Montgomery, et al.
Published: (2023)
by: Gole, Montgomery, et al.
Published: (2023)
AIDetx: a compression-based method for identification of machine-learning generated text
by: Almeida, Leonardo, et al.
Published: (2024)
by: Almeida, Leonardo, et al.
Published: (2024)
Universal Neurons in GPT2 Language Models
by: Gurnee, Wes, et al.
Published: (2024)
by: Gurnee, Wes, et al.
Published: (2024)
Machine learning methods fail to provide cohesive atheoretical construction of personality traits from semantic embeddings
by: Bouguettaya, Ayoub, et al.
Published: (2025)
by: Bouguettaya, Ayoub, et al.
Published: (2025)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
by: Grishina, Ekaterina, et al.
Published: (2025)
by: Grishina, Ekaterina, et al.
Published: (2025)
Is Next Token Prediction Sufficient for GPT? Exploration on Code Logic Comprehension
by: Qi, Mengnan, et al.
Published: (2024)
by: Qi, Mengnan, et al.
Published: (2024)
Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
by: Hackl, Veronika, et al.
Published: (2023)
by: Hackl, Veronika, et al.
Published: (2023)
ChatGPT for automated grading of short answer questions in mechanical ventilation
by: Jade, Tejas, et al.
Published: (2025)
by: Jade, Tejas, et al.
Published: (2025)
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
FoldGPT: Simple and Effective Large Language Model Compression Scheme
by: Liu, Songwei, et al.
Published: (2024)
by: Liu, Songwei, et al.
Published: (2024)
Similar Items
-
Challenges and Considerations in Annotating Legal Data: A Comprehensive Overview
by: Darji, Harshil, et al.
Published: (2024) -
Compressed Concatenation of Small Embedding Models
by: Ayad, Mohamed Ayoub Ben, et al.
Published: (2025) -
Technical Report: Impact of Position Bias on Language Models in Token Classification
by: Amor, Mehdi Ben, et al.
Published: (2023) -
Computational Approaches to the Detection of Lesser-Known Rhetorical Figures: A Systematic Survey and Research Challenges
by: Kühn, Ramona, et al.
Published: (2024) -
Enhancing Rhetorical Figure Annotation: An Ontology-Based Web Application with RAG Integration
by: Kühn, Ramona, et al.
Published: (2024)