Skip to content

Epoch model arguments can be overridden by equivalent local CLI spellings #1911

Open @vitaly-andr opened 2026-10-03 11:58 UTC 0 comments Updated 2026-10-03 11:58 UTC

Problem

MergeModelArgs compares raw CLI tokens when it gives epoch model arguments precedence over a host's local arguments. These spellings name the same vLLM option but both survive today:

epoch: ["--max-model-len", "240000"]
local: ["--max-model-len=8192"]
merged: ["--max-model-len", "240000", "--max-model-len=8192"]

vLLM uses the later scalar value. The same mismatch exists when the epoch uses = and the local setting uses a space, underscores, or an accepted long abbreviation such as --max-model. Short aliases take a different path: the broker silently drops them. This affects other model options as well as the context limit. It also becomes relevant to #1763, which bases the inference reservation cap on the epoch snapshot's context limit.

How do you know this is a real problem?

The raw-string comparison is in decentralized-api/broker/broker.go. A broker regression test using the argument list above fails on the head of #1763 because the local option remains in the deployment arguments. This is a source-level reproduction; I have not launched the pinned vLLM image with these arguments.

The issue predates #1763. That PR changes the reservation path but does not change MergeModelArgs.

Expected behavior

When epoch and local arguments refer to the same vLLM option, the epoch value should control the deployment. Unrelated local options should retain their order. MLNode's own parsing of TP/PP and context arguments should agree with the values it passes to vLLM. A deployment whose effective arguments change may need to restart, so the fix should document that impact.


🔄 Auto-synced from Issue #1911 every hour.