Skip to content

[Ops][Feature] Add prefill abort side-channel support for disaggregated prefill - #14716

Open
UpDown9 wants to merge 1 commit into
vllm-project:rfc/vllm_cannfrom
UpDown9:max-token
Open

[Ops][Feature] Add prefill abort side-channel support for disaggregated prefill#14716
UpDown9 wants to merge 1 commit into
vllm-project:rfc/vllm_cannfrom
UpDown9:max-token

Conversation

@UpDown9

@UpDown9 UpDown9 commented Aug 21, 2026

Copy link
Copy Markdown

What this PR does / why we need it?

This PR introduces a prefill abort side-channel mechanism for disaggregated prefill. It allows the load balance proxy to notify prefillers to abort and force-release resources when a decoder stream is cancelled or encounters an error. This is achieved by introducing a new ABORT_REQUEST_MSG message type in the Mooncake_hybrid_connector and scheduling abort requests via ZeroMQ (zmq).

Feedback on the changes:

  • In _send_prefill_abort_sync, creating and terminating a new zmq.Context on every abort request is inefficient and can cause resource leaks. It is recommended to use zmq.Context.instance() instead.
  • In assign_instances, kv_transfer_params should be safely defaulted to prevent potential AttributeError or TypeError if it is None.

Does this PR introduce any user-facing change?

No, this is an internal optimization and error-handling mechanism for disaggregated prefill.

How was this patch tested?

CI passed with existing tests.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces an abort API for P-instances (prefillers) to allow the decoder to proactively cancel requests. By implementing a ZMQ-based side-channel, the proxy server can now notify prefillers to release resources associated with specific request IDs when an error occurs or a task is cancelled. This change improves system robustness and resource efficiency in disaggregated prefill scenarios.

Highlights

  • P-instance Abort API: Introduced a new abort mechanism allowing the decoder to signal prefillers to cancel specific requests, improving resource management.
  • ZMQ Communication: Implemented ZMQ-based messaging to facilitate communication between the proxy server and prefillers for request abortion.
  • Error Handling: Enhanced error handling in the proxy server to ensure that aborted or failed requests are properly cleaned up and acknowledged.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Add prefill abort side-channel support for disaggregated prefill

Suggested PR Summary:

### What this PR does / why we need it?

This PR introduces a prefill abort side-channel mechanism for disaggregated prefill. It allows the load balance proxy to notify prefillers to abort and force-release resources when a decoder stream is cancelled or encounters an error. This is achieved by introducing a new `ABORT_REQUEST_MSG` message type in the `Mooncake_hybrid_connector` and scheduling abort requests via ZeroMQ (`zmq`).

Feedback on the changes:
- In `_send_prefill_abort_sync`, creating and terminating a new `zmq.Context` on every abort request is inefficient and can cause resource leaks. It is recommended to use `zmq.Context.instance()` instead.
- In `assign_instances`, `kv_transfer_params` should be safely defaulted to prevent potential `AttributeError` or `TypeError` if it is `None`.

### Does this PR introduce _any_ user-facing change?

No, this is an internal optimization and error-handling mechanism for disaggregated prefill.

### How was this patch tested?

CI passed with existing tests.

Comment thread examples/disaggregated_prefill_v1/load_balance_proxy_server_example.py Outdated
@UpDown9 UpDown9 changed the title add P instance abort api [Ops][Feature] Add prefill abort side-channel support for disaggregated prefill Aug 21, 2026
@UpDown9
UpDown9 force-pushed the max-token branch 2 times, most recently from 86c5bb0 to 0e128ba Compare August 21, 2026 06:39
Signed-off-by: xujiuxu9 <xujiuxu1@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant