Repository navigation
Conversation
…P8 (Ornith) Some compressed-tensors checkpoints (e.g. Ornith-1.5) store FP8 weights in the float-quantized format with a per-output-channel .weight_scale (BF16, shape [N, 1]) instead of the blocked .weight_scale_inv ([K//128, N//128]). The converter rejected them outright (pack-quantized-only assert) and the FP8 path did not recognize the .weight_scale key. - converter: accept float-quantized 8-bit float compressed-tensors groups (group_0 / config_group_0), map to FP8Format(block_out=1) for per-channel and block_out=128 for blocked - FP8Format: recognize .weight_scale, accept per-channel [N, 1] scales, add post_process hook that expands per-channel scales to the K-grouped kernel layout [K//128, N] (identical rows; weight bytes unchanged) - WeightFormat: identity post_process hook on the base class The C++ side already supports per-channel FP8 (groupwise 128x1 e4m3, InternLM#4943); this only closes the python loading gap.
…_process normalize() transposes 2-D tensors to TM layout before post_process runs, so the per-channel weight_scale [N, 1] reaches the hook as [1, N]. Match the transposed shape when expanding to [K//128, N]; guard also rejects any other layout instead of crashing on expand.
…ors float-quantized checkpoints Ornith-style checkpoints carry quant_method=compressed-tensors with float-quantized FP8 weights. The production startup script passes an explicit --model-format fp8, which previously hit the strict 'user input == quant_method' assert. fp8 and compressed-tensors float-quantized are the same underlying format, so accept the override; the float-quantized branch still validates 8-bit float weights.
CI lint/unit_test fixes for the per-channel FP8 PR: - converter.py: wrap the _build_quantized_formats signature to respect the 120-char line-length limit (ruff E501). - test_fp8_per_channel.py: sort imports (ruff I001) and fix the two test cases that call FP8Format.post_process / dequant directly to pass the normalize()-transposed per-channel scale [1, N] (not the raw [N, 1]); this matches the production path where normalize() transposes 2-D tensors to TM layout before post_process runs. Verified in a real-CUDA container (lmdeploy 0.18.0, cu128): all 6 tests pass; regression A/B (patched vs pristine) on the existing linear tests is identical (1 pre-existing skip-path failure in both), so the patch breaks nothing.
CI lint's docformatter hook (v1.7.7, --wrap-descriptions 120) requires a blank line after the summary sentence of the base-class post_process docstring I added. Split it into summary + body so pre-commit passes.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
TurboMind cannot load per-channel (channel-quantized) FP8 compressed-tensors checkpoints such as
Ornith-1.5-35B-A3B-FP8:quant_method: compressed-tensors,format: float-quantizedfloatweights (F8_E4M3),strategy: channel,group_size: null*.weight_scale(BF16, shape[N, 1]), group keyconfig_group_0The converter hard-asserted
pack-quantized+ 4-bit int for compressed-tensors, so these checkpoints were rejected.Changes (python-only; C++ already supports it)
Mainline already has the FP8 compute path for pre-SM89 GPUs (#4871) and groupwise
128x1e4m3 kernels (#4943). This PR only closes the python-side gap:converter.py: acceptfloat-quantizedcompressed-tensors with 8-bit float weights (per-channel ->FP8Format(block_out=1), 128x128 blocked ->block_out=128); allow explicit--model-format fp8for such checkpoints; group-key fallbackconfig_group_0/group_0weight_format.py:FP8Formataccepts.weight_scalesuffix and per-channel[N, 1]scales; newpost_processhook expands per-channel scales into the kernel's K-grouped layout[K//128, N](identity for blocked, so no behavior change for existing fp8 checkpoints)builders/linear.py:_build_linearinvokes the format'spost_processtests/turbomind/linear/test_fp8_per_channel.py: new unit testsValidation (2x RTX 2080Ti, sm75, CUDA 12.8)
Ornith-1.5-35B-A3B-FP8e2e: converter parse -> per-channel scale expansion -> e4m3 groupwise GEMM plan -> serve + chat all pass