Skip to content

Add chunking strategy for fp8_paged_mqa_logits - #398

Draft
Xia-Weiwen wants to merge 1 commit into
mainfrom
fp8_mqa_logits_oom
Draft

Add chunking strategy for fp8_paged_mqa_logits#398
Xia-Weiwen wants to merge 1 commit into
mainfrom
fp8_mqa_logits_oom

Conversation

@Xia-Weiwen

Copy link
Copy Markdown
Collaborator

export SGL_KERNEL_FP8_PAGED_MQA_CHUNK_MB=512 (default is 512)

@polisettyvarma polisettyvarma left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

helps in OOM and which model, right ?

@Xia-Weiwen

Copy link
Copy Markdown
Collaborator Author

helps in OOM and which model, right ?

It can probably resolve the OOM issue of DeepSeek. Jianan is going to have a try.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants