Skip to main navigation Skip to search Skip to main content

TransKV: A Data-Driven Pruning Method for Large Foundation Models

  • Guangning Xu
  • , Fanxu Meng*
  • , Ruikai Zhou
  • , Kwok Po Ng
  • , Wenjie Pei
  • , Muhan Zhang
  • *Corresponding author for this work

Research output: Chapter in book/report/conference proceedingConference proceedingpeer-review

Abstract

Large Foundation Models (LFMs) have achieved remarkable success across various domains, yet their enormous parameter counts severely hinder deployment on resource-constrained devices. Recent pruning methods leverage low-rank structures in attention mechanisms, with SOTA techniques like CLOVER and Olica targeting redundancies in weight pairs (e.g., W^ i _Q (W^ i _K)^\top, i-th attention head) to enable more aggressive compression than single-matrix pruning. However, these static approaches overlook data-driven interactions, as singular value distributions of static weight pairs differ significantly from data-driven counterparts (e.g., X W^ i _Q (W^ i _K)^\top X^\top), often yielding suboptimal performance. To address this, we introduce TransKV, a data-driven pruning framework that derives a projection matrix from keys (X W^ i _K) and values (X W^ i _V), seamlessly integrating it into weight-pair computations for precise low-rank truncation. TransKV natively supports core LFM components, including Multi-Head Attention (MHA), Grouped Query Attention (GQA), RMSNorm, and Rotary Position Embeddings (RoPE), ensuring compatibility with contemporary LFMs. Extensive experiments on diverse benchmarks across multiple domains validate TransKV's superiority, outperforming SOTA methods on various LFMs (including DeepSeek-OCR) at equivalent compression ratios.
Original languageEnglish
Title of host publicationProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings 2026
PublisherIEEE
Pages2451-2461
Number of pages11
Publication statusPublished - Jun 2026
Event2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2026 - Denver, United States
Duration: 3 Jun 20267 Jun 2026
https://cvpr.thecvf.com/Conferences/2026
https://cvpr.thecvf.com/virtual/2026/index.html
https://cvpr.thecvf.com/virtual/2026/papers.html
https://media.eventhosts.cc/Conferences/CVPR2026/CVPR_main_conf_2026_15.pdf
https://openaccess.thecvf.com/CVPR2026

Conference

Conference2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2026
Country/TerritoryUnited States
CityDenver
Period3/06/267/06/26
Internet address

Fingerprint

Dive into the research topics of 'TransKV: A Data-Driven Pruning Method for Large Foundation Models'. Together they form a unique fingerprint.

Cite this