-
Notifications
You must be signed in to change notification settings - Fork 837
cookbook(diffusers): add FP8 quantized checkpoint support #304
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
ConstBob
wants to merge
3
commits into
NVIDIA:main
Choose a base branch
from
ConstBob:feat/diffusers-fp8
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from 2 commits
Commits
Show all changes
3 commits
Select commit
Hold shift + click to select a range
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Not sure what the final model-card/repository structure will be, but we need important fix for the current
nvidia/Cosmos3-Experimentalsubfolders to fix Distilled FP8. Because Distilled models can only be used with Modular pipeline, and how weights and quantization state get loaded for regular and modular pipelines differ.The regular loader joins the component path first, while the modular loader keeps the subfolder separate, which ModelOpt ignores during state lookup.
Fixing it here is safer than changing generic diffusers loading or depending fixing in ModelOpt, which will only land in 0.46+ version and will be incompatible with 0.44 used for quantization already.
So the fix is -- load the transformer from its full path, pin it, then load the remaining components normally:
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
There might be a better solution, so feel free to handle it other way. This is also something to fix for the HF model cards examples, although this is dependent on the final HF repo structure.
Some helper that I used to verify correctness:
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Thanks for the review and suggestions! I implemented the Distilled FP8 preload approach you suggested (full-path transformer load →
update_components→load_components, plus a small helper to materialize Hub checkpoints locally) and added yourverify_fp8helper: f4878b6Below testing has been conducted: On A100-80GB cluster (
nvidia-modelopt==0.44.0), I loaded Experimental Distilled T2I 4-Step FP8 through that path and ranverify_fp8("Cosmos3-Super-Text2Image-4Step"); ModelOpt restoredtransformer/modelopt_state.pthand verification passed (weights=896, quantizers=2709, meta=0).There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Thanks! Let's wait till checkpoints are released in the final HF model repos to update paths in the cookbook before merging it