An Empirical Investigation of Matrix Factorization Methods for Pre-trained Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Gupta, Ashim, Saravani, Sina Mahdipour, Sadayappan, P., Srikumar, Vivek |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
di: Asudeh, Omid, et al.
Pubblicazione: (2025)
di: Asudeh, Omid, et al.
Pubblicazione: (2025)
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
di: Gupta, Ashim, et al.
Pubblicazione: (2025)
State Space Models are Strong Text Rerankers
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Enhancing Question Answering on Charts Through Effective Pre-training Tasks
di: Gupta, Ashim, et al.
Pubblicazione: (2024)
di: Gupta, Ashim, et al.
Pubblicazione: (2024)
Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
di: Gupta, Ashim, et al.
Pubblicazione: (2023)
di: Gupta, Ashim, et al.
Pubblicazione: (2023)
Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion
di: Mehta, Maitrey, et al.
Pubblicazione: (2026)
di: Mehta, Maitrey, et al.
Pubblicazione: (2026)
Unequal Voices: How LLMs Construct Constrained Queer Narratives
di: Ghosal, Atreya, et al.
Pubblicazione: (2025)
di: Ghosal, Atreya, et al.
Pubblicazione: (2025)
Reinforcing Code Generation: Improving Text-to-SQL with Execution-Based Learning
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
LLM-Symbolic Integration for Robust Temporal Tabular Reasoning
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharv, et al.
Pubblicazione: (2025)
InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis
di: Bentham, Oliver, et al.
Pubblicazione: (2026)
di: Bentham, Oliver, et al.
Pubblicazione: (2026)
Understanding the Logic of Direct Preference Alignment through Logic
di: Richardson, Kyle, et al.
Pubblicazione: (2024)
di: Richardson, Kyle, et al.
Pubblicazione: (2024)
Promptly Predicting Structures: The Return of Inference
di: Mehta, Maitrey, et al.
Pubblicazione: (2024)
di: Mehta, Maitrey, et al.
Pubblicazione: (2024)
Hybrid Dialogue State Tracking for Persian Chatbots: A Language Model-Based Approach
di: Aghabagher, Samin Mahdipour, et al.
Pubblicazione: (2025)
di: Aghabagher, Samin Mahdipour, et al.
Pubblicazione: (2025)
In-Context Example Ordering Guided by Label Distributions
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
di: Xu, Zhichao, et al.
Pubblicazione: (2024)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
Distillation versus Contrastive Learning: How to Train Your Rerankers
di: Xu, Zhichao, et al.
Pubblicazione: (2025)
di: Xu, Zhichao, et al.
Pubblicazione: (2025)
PART: Pre-trained Authorship Representation Transformer
di: Huertas-Tato, Javier, et al.
Pubblicazione: (2022)
di: Huertas-Tato, Javier, et al.
Pubblicazione: (2022)
Pre-trained Language Models for Keyphrase Generation: A Thorough Empirical Study
di: Wu, Di, et al.
Pubblicazione: (2022)
di: Wu, Di, et al.
Pubblicazione: (2022)
On Initializing Transformers with Pre-trained Embeddings
di: Kim, Ha Young, et al.
Pubblicazione: (2024)
di: Kim, Ha Young, et al.
Pubblicazione: (2024)
BERTtime Stories: Investigating the Role of Synthetic Story Data in Language Pre-training
di: Theodoropoulos, Nikitas, et al.
Pubblicazione: (2024)
di: Theodoropoulos, Nikitas, et al.
Pubblicazione: (2024)
HQFS: Hybrid Quantum Classical Financial Security with VQC Forecasting, QUBO Annealing, and Audit-Ready Post-Quantum Signing
di: Nayak, Srikumar
Pubblicazione: (2026)
di: Nayak, Srikumar
Pubblicazione: (2026)
Named Entity Recognition for Payment Data Using NLP
di: Nayak, Srikumar
Pubblicazione: (2026)
di: Nayak, Srikumar
Pubblicazione: (2026)
Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
di: Lindemann, Matthias, et al.
Pubblicazione: (2024)
di: Lindemann, Matthias, et al.
Pubblicazione: (2024)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
di: Dawkins, Hillary, et al.
Pubblicazione: (2024)
di: Dawkins, Hillary, et al.
Pubblicazione: (2024)
Investigating Data Contamination for Pre-training Language Models
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
The Dark Side of the Language: Pre-trained Transformers in the DarkNet
di: Ranaldi, Leonardo, et al.
Pubblicazione: (2022)
di: Ranaldi, Leonardo, et al.
Pubblicazione: (2022)
RLShield: Practical Multi-Agent RL for Financial Cyber Defense with Attack-Surface MDPs and Real-Time Response Orchestration
di: Nayak, Srikumar
Pubblicazione: (2026)
di: Nayak, Srikumar
Pubblicazione: (2026)
Rethinking Transformer-based Multi-document Summarization: An Empirical Investigation
di: Ma, Congbo, et al.
Pubblicazione: (2024)
di: Ma, Congbo, et al.
Pubblicazione: (2024)
Pre-trained Transformer-Based Approach for Arabic Question Answering : A Comparative Study
di: Alsubhi, Kholoud, et al.
Pubblicazione: (2021)
di: Alsubhi, Kholoud, et al.
Pubblicazione: (2021)
A Comparative Study of Pre-training and Self-training
di: Wang, Yiheng, et al.
Pubblicazione: (2024)
di: Wang, Yiheng, et al.
Pubblicazione: (2024)
On Predicting the Post-training Potential of Pre-trained LLMs
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2026)
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
di: Luo, Zheheng, et al.
Pubblicazione: (2024)
TransGPT: Multi-modal Generative Pre-trained Transformer for Transportation
di: Wang, Peng, et al.
Pubblicazione: (2024)
di: Wang, Peng, et al.
Pubblicazione: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
di: Labrak, Yanis, et al.
Pubblicazione: (2025)
di: Labrak, Yanis, et al.
Pubblicazione: (2025)
Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models
di: Zhuang, Xinlin, et al.
Pubblicazione: (2025)
di: Zhuang, Xinlin, et al.
Pubblicazione: (2025)
Multiple-Debias: A Full-process Debiasing Method for Multilingual Pre-trained Language Models
di: Liang, Haoyu, et al.
Pubblicazione: (2026)
di: Liang, Haoyu, et al.
Pubblicazione: (2026)
Domain Pre-training Impact on Representations
di: Gonzalez-Gutierrez, Cesar, et al.
Pubblicazione: (2025)
di: Gonzalez-Gutierrez, Cesar, et al.
Pubblicazione: (2025)
TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting
di: Cao, Defu, et al.
Pubblicazione: (2023)
di: Cao, Defu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
di: Asudeh, Omid, et al.
Pubblicazione: (2025) -
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
di: Gupta, Ashim, et al.
Pubblicazione: (2025) -
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
di: Gupta, Ashim, et al.
Pubblicazione: (2025) -
State Space Models are Strong Text Rerankers
di: Xu, Zhichao, et al.
Pubblicazione: (2024) -
Enhancing Question Answering on Charts Through Effective Pre-training Tasks
di: Gupta, Ashim, et al.
Pubblicazione: (2024)