Skip to content

fix: proxy get_node_url raises TypeError when no node matches role - #5013

Open
Evolian-o wants to merge 1 commit into
InternLM:mainfrom
Evolian-o:fix-proxy-get-node-url-unpack
Open

Evolian-o wants to merge 1 commit into
InternLM:mainfrom
Evolian-o:fix-proxy-get-node-url-unpack

Conversation

@Evolian-o

Copy link
Copy Markdown

Summary

NodeManager.get_node_url crashes with TypeError: cannot unpack non-iterable NoneType object when no node matches the requested model + role, returning HTTP 500 instead of an "unavailable model" response.

Root cause

get_matched_urls() (nested in get_node_url) returns None when there is no matching node:

all_matched_urls = urls_with_speeds + urls_without_speeds
if len(all_matched_urls) == 0:
    return None

But both the RANDOM and MIN_EXPECTED_LATENCY branches unpack the return value directly:

all_matched_urls, all_the_speeds = get_matched_urls()

so the None raises TypeError before the if len(all_matched_urls) == 0: return None guard is ever reached (it is dead code).

When this happens

In DistServe (PD-disaggregation) mode, a model can be registered on a node of one role while the engine of the requested role is unavailable (e.g. the Prefill engine is down or removed by the heartbeat check). check_request_model still passes because the model is in the global model list, but get_node_url(model, EngineRole.Prefill) finds no Prefill node and crashes instead of returning an "unavailable model" response.

Fix

Make get_matched_urls always return a (urls, speeds) pair (an empty pair instead of None), so the existing len(...) == 0 guard becomes reachable and get_node_url returns None as intended.

Tests

Added tests/test_lmdeploy/serve/test_proxy.py covering all three routing strategies for both "no node matches" (returns None) and "node matches" (returns the URL).

get_matched_urls() returns None when no node matches the requested model
and role, but the RANDOM and MIN_EXPECTED_LATENCY branches unpack its
return value directly, raising "TypeError: cannot unpack non-iterable
NoneType object" and returning HTTP 500 instead of an "unavailable model"
response. Return an empty pair instead so the existing len() == 0 guard
is reachable.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant