ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yongkang, Wang, Zijing, Zhao, Mengjie, Nie, Ercong, Wang, Mingyang, Li, Qian, Ren, Feiliang, Feng, Shi, Wang, Daling, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2025)
by: Liu, Yongkang, et al.
Published: (2025)
High-Rank Structured Modulation for Parameter-Efficient Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging
by: Wang, Zijing, et al.
Published: (2026)
by: Wang, Zijing, et al.
Published: (2026)
PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs
by: Wang, Zijing, et al.
Published: (2026)
by: Wang, Zijing, et al.
Published: (2026)
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset
by: Liu, Yongkang, et al.
Published: (2026)
by: Liu, Yongkang, et al.
Published: (2026)
A Unified Data Augmentation Framework for Low-Resource Multi-Domain Dialogue Generation
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
by: Nie, Ercong, et al.
Published: (2025)
by: Nie, Ercong, et al.
Published: (2025)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
by: Liu, Yongkang, et al.
Published: (2023)
by: Liu, Yongkang, et al.
Published: (2023)
ChatZero:Zero-shot Cross-Lingual Dialogue Generation via Pseudo-Target Language
by: Liu, Yongkang, et al.
Published: (2024)
by: Liu, Yongkang, et al.
Published: (2024)
BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context Learning
by: Nie, Ercong, et al.
Published: (2024)
by: Nie, Ercong, et al.
Published: (2024)
Why Do More Experts Fail? A Theoretical Analysis of Model Merging
by: Wang, Zijing, et al.
Published: (2025)
by: Wang, Zijing, et al.
Published: (2025)
Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network
by: Yuan, Shuzhou, et al.
Published: (2024)
by: Yuan, Shuzhou, et al.
Published: (2024)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
Tracing Multilingual Factual Knowledge Acquisition in Pretraining
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
by: Liu, Yihong, et al.
Published: (2023)
by: Liu, Yihong, et al.
Published: (2023)
Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence
by: He, Linyang, et al.
Published: (2024)
by: He, Linyang, et al.
Published: (2024)
Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
On the Entity-Level Alignment in Crosslingual Consistency
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
by: Nie, Ercong, et al.
Published: (2024)
by: Nie, Ercong, et al.
Published: (2024)
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks
by: Ma, Bolei, et al.
Published: (2024)
by: Ma, Bolei, et al.
Published: (2024)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
LLM in the Loop: Creating the ParaDeHate Dataset for Hate Speech Detoxification
by: Yuan, Shuzhou, et al.
Published: (2025)
by: Yuan, Shuzhou, et al.
Published: (2025)
MEKiT: Multi-source Heterogeneous Knowledge Injection Method via Instruction Tuning for Emotion-Cause Pair Extraction
by: Mu, Shiyi, et al.
Published: (2025)
by: Mu, Shiyi, et al.
Published: (2025)
NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis
by: Wang, Mingyang, et al.
Published: (2023)
by: Wang, Mingyang, et al.
Published: (2023)
Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization
by: Wang, Mingyang, et al.
Published: (2024)
by: Wang, Mingyang, et al.
Published: (2024)
LangSAMP: Language-Script Aware Multilingual Pretraining
by: Liu, Yihong, et al.
Published: (2024)
by: Liu, Yihong, et al.
Published: (2024)
Learn it or Leave it: Module Composition and Pruning for Continual Learning
by: Wang, Mingyang, et al.
Published: (2024)
by: Wang, Mingyang, et al.
Published: (2024)
Rehearsal-Free Modular and Compositional Continual Learning for Language Models
by: Wang, Mingyang, et al.
Published: (2024)
by: Wang, Mingyang, et al.
Published: (2024)
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
From Knowledge to Noise: CTIM-Rover and the Pitfalls of Episodic Memory in Software Engineering Agents
by: Lindenbauer, Tobias, et al.
Published: (2025)
by: Lindenbauer, Tobias, et al.
Published: (2025)
RevFFN: Memory-Efficient Full-Parameter Fine-Tuning of Mixture-of-Experts LLMs with Reversible Blocks
by: Liu, Ningyuan, et al.
Published: (2025)
by: Liu, Ningyuan, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning with Discrete Fourier Transform
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
by: Yan, Sikuan, et al.
Published: (2025)
by: Yan, Sikuan, et al.
Published: (2025)
Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal Causes
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
Bring Your Own Knowledge: A Survey of Methods for LLM Knowledge Expansion
by: Wang, Mingyang, et al.
Published: (2025)
by: Wang, Mingyang, et al.
Published: (2025)
LongForm: Effective Instruction Tuning with Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2023)
by: Köksal, Abdullatif, et al.
Published: (2023)
S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain
by: Zhang, Baoquan, et al.
Published: (2026)
by: Zhang, Baoquan, et al.
Published: (2026)
Similar Items
-
Look Within or Look Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2025) -
High-Rank Structured Modulation for Parameter-Efficient Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026) -
SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning
by: Liu, Yongkang, et al.
Published: (2026) -
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy
by: Liu, Yongkang, et al.
Published: (2024) -
DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging
by: Wang, Zijing, et al.
Published: (2026)