Abstract
Training efficiency in large-scale models is typically assessed through memory consumption, training time, and model performance. Current methods often exhibit trade-offs among these metrics, as optimizing one generally degrades at least one of the others. Addressing this trade-off remains a central challenge in algorithm design. While GaLore enables memory-efficient training by updating gradients in a low-rank subspace, it incurs a comparable extra training time cost due to the Singular Value Decomposition(SVD) process on gradients. In this paper, we propose Lotus, a method that resolves this trade-off by simply modifying the projection process. We propose a criterion that quantifies the displacement of the unit gradient to enable efficient transitions between low-rank gradient subspaces. Experimental results indicate that Lotus is the most efficient method, achieving a 30% reduction in training time and a 40% decrease in memory consumption for gradient and optimizer states. Additionally, it outperforms the baseline method in both pre-training and fine-tuning tasks.
| Original language | English |
|---|---|
| Title of host publication | ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) |
| Place of Publication | Barcelona |
| Publisher | IEEE |
| Pages | 4696-4700 |
| Number of pages | 5 |
| ISBN (Electronic) | 9798331567019 |
| ISBN (Print) | 9798331567026 |
| DOIs | |
| Publication status | Published - 3 May 2026 |
| Event | 2026 IEEE International Conference on Acoustics, Speech and Signal Processing - Centre de Convencions Internacional de Barcelona, Barcelona, Spain Duration: 3 May 2026 → 8 May 2026 https://2026.ieeeicassp.org/ (Conference website) https://2026.ieeeicassp.org/technical-program/ (Conference program schedule) https://ieeexplore.ieee.org/xpl/conhome/11460365/proceeding (Conference proceeding) |
Publication series
| Name | IEEE International Conference on Acoustics, Speech and Signal Processing |
|---|---|
| Publisher | IEEE |
| ISSN (Print) | 1520-6149 |
| ISSN (Electronic) | 2379-190X |
Conference
| Conference | 2026 IEEE International Conference on Acoustics, Speech and Signal Processing |
|---|---|
| Abbreviated title | ICASSP 2026 |
| Country/Territory | Spain |
| City | Barcelona |
| Period | 3/05/26 → 8/05/26 |
| Internet address |
|
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
User-Defined Keywords
- Efficient Training
- Pre-training
- Fine-tuning
- Large Language Model
Fingerprint
Dive into the research topics of 'Lotus: Efficient LLM Training By Randomized Low-Rank Gradient Projection with Adaptive Subspace Switching'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver