Skip to content

Add configurable per-layer LoRA ranks (Related to Issue #2335) - #2336

Open
kor80 wants to merge 1 commit into
Lightning-AI:mainfrom
kor80:feat/layer-wise-lora
Open

kor80 wants to merge 1 commit into
Lightning-AI:mainfrom
kor80:feat/layer-wise-lora

Conversation

@kor80

@kor80 kor80 commented Sep 25, 2026 •

Copy link
Copy Markdown

Overview

This pull request introduces configurable static per-layer LoRA ranks, allowing users to assign different ranks to individual transformer blocks while preserving the existing global rank behaviour when no layer-specific configuration is provided.

Motivations

Static allocation makes it possible to assign different parameter budgets to individual transformer blocks without introducing additional training-time machinery. The allocation is explicitly defined before training, making configurations simple and reproducible.

Changes

  • Add optional lora_r_by_layer configuration with input validation;
  • Support block-specific ranks across attention, MLP and MoE modules;
  • Retain the global rank for unspecified blocks and the language modeling head;
  • Integrate the new option into the fine-tuning CLI;
  • Support serialisation and restoration of layer-specific ranks through checkpoint metadata;
  • Add test and usage documentation.

Testing

Added tests covering rank allocation, configuration validation, backward compatibility, checkpointing sanity check and training updates.

Limitations

  • Rank allocation must be specified manually rather than learned or adapted during training.
  • This implementation provides flexible parameter allocation but does not automatically identify which layers benefit the most from adaptation.

Closes #2335

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Configurable per-layer LoRA ranks

1 participant