Modelplane Modelplane docs

Qwen/Qwen3-8B

An 8.2B dense chat model on a single NVIDIA L4.

View on Hugging Face

An 8.2B dense chat model on a single NVIDIA L4. The smallest recipe: one Standalone engine, no cache, weights pulled straight from Hugging Face.

This recipe was run end to end; the InferenceClass and ModelDeployment are the exact manifests from that run. Apply the platform side first, then the ML side.

Validated deployments

Dense 8B 16,384 ctx vLLM
Cloud
Google Cloud Nebius
GPU
L4 24G A100 40/80G H100 80G H200 141G 1× per node
Serving mode
Standalone LeaderWorker PrefillDecode
Precision
Engine
vLLM SGLang llama.cpp
Image
vllm/vllm-openai:v0.23.0
Features
Speculative decoding for low latency and small batch sizes

Platform

inference-class.yaml
# InferenceClass for the L4 shape, validated serving Qwen3-8B on EKS.
#
# One NVIDIA L4 on an EKS g6.xlarge. The single GPU is a claim: DRA device;
# the scheduler matches a ModelDeployment's nodeSelector against its declared
# capacity and DRA binds it to the serving pod.
apiVersion: modelplane.ai/v1alpha1
kind: InferenceClass
metadata:
  name: eks-l4-1x-g6
spec:
  description: "EKS g6.xlarge, 1x NVIDIA L4"
  provisioning:
    provider: EKS
    eks:
      instanceType: g6.xlarge
      diskSizeGb: 100
      accelerator:
        type: nvidia-l4
        count: 1
  devices:
  - name: gpu
    claim: DRA
    driver: gpu.nvidia.com
    deviceClassName: gpu.nvidia.com
    count: 1
    attributes:
      architecture: { string: Ada Lovelace }
    capacity:
      # The L4's real usable VRAM as the NVIDIA DRA driver reports it, not the
      # nominal 24GB.
      memory: { value: "23034Mi" }
inference-cluster.yaml
# An EKS InferenceCluster with one L4 node pool, labeled for the
# ModelDeployment's clusterSelector to target.
apiVersion: modelplane.ai/v1alpha1
kind: InferenceCluster
metadata:
  name: eks-l4
  labels:
    modelplane.ai/region: us
spec:
  cluster:
    source: EKS
    eks:
      region: us-west-2
  nodePools:
  - name: gpu-l4
    className: eks-l4-1x-g6
    nodeCount: 1
    minNodeCount: 1
    maxNodeCount: 1
    zones:
    - us-west-2a

Deployment

model-deployment.yaml
# Qwen3-8B served on a single NVIDIA L4, validated end to end on EKS.
#
# An 8.2B dense model is a single Standalone engine: one self-contained vLLM
# pod, no ModelCache, weights pulled straight from Hugging Face. The flags carry
# real meaning beyond fit:
#
#   --tool-call-parser=hermes        the parser for Qwen3 dense (qwen3_xml is
#                                    for Qwen3-Coder, not this model). Qwen3's
#                                    tool-use template ships in the tokenizer,
#                                    so no --chat-template is needed.
#   --reasoning-parser=qwen3 with
#   --default-chat-template-kwargs   turns thinking off. Qwen3 thinks by
#                                    default, burying a one-line answer under a
#                                    <think> block and forbidding greedy decode.
#   --max-model-len / --gpu-memory-utilization  L4 fit, not correctness.
#
# No --port or --host: Modelplane's routing expects the engine on its default
# :8000 with a /health probe, and passes args through verbatim.
apiVersion: modelplane.ai/v1alpha1
kind: ModelDeployment
metadata:
  name: qwen3-8b
  namespace: ml-team
spec:
  replicas: 1
  template:
    spec:
      clusterSelector:
        matchLabels:
          modelplane.ai/region: us
      engines:
      - name: qwen3-8b
        members:
        - role: Standalone
          nodeSelector:
            devices:
            - name: gpu
              count: 1
              selectors:
              - cel: |
                  device.capacity["gpu.nvidia.com"].memory.compareTo(quantity("20Gi")) >= 0
          template:
            spec:
              containers:
              - name: engine
                image: vllm/vllm-openai:v0.23.0
                args:
                - "--model=Qwen/Qwen3-8B"
                - "--served-model-name=qwen"
                - "--max-model-len=16384"
                - "--gpu-memory-utilization=0.92"
                - "--reasoning-parser=qwen3"
                - "--default-chat-template-kwargs={\"enable_thinking\": false}"
                - "--enable-auto-tool-choice"
                - "--tool-call-parser=hermes"
model-service.yaml
# Exposes the qwen3-8b deployment's endpoints as a single OpenAI-compatible URL.
# Modelplane labels each composed ModelEndpoint with the deployment name, so this
# selector reaches every replica. Read the public address from status.address:
#   kubectl get ms qwen3-8b -n ml-team -o jsonpath='{.status.address}'
apiVersion: modelplane.ai/v1alpha1
kind: ModelService
metadata:
  name: qwen3-8b
  namespace: ml-team
spec:
  endpoints:
  - selector:
      matchLabels:
        modelplane.ai/deployment: qwen3-8b

Speculative decoding

The same model and platform also serve with n-gram (prompt-lookup) speculative decoding, which proposes tokens by matching the prompt and so needs no draft model or second set of weights. On copy-heavy output, editing a pasted code block where most output tokens are copied from the prompt, it roughly doubles decode throughput and halves the time per output token:

MetricWithout speculationWith n-gram speculation
Output token throughput (tok/s)16.1039.01
Mean TPOT (ms/token)60.2024.21

Measured on a single L4 (vllm/vllm-openai:v0.23.0, Qwen3-8B, 30 copy-heavy prompts at concurrency 1) against the same model without --speculative-config; the speculative run accepted 65% of drafted tokens, a mean acceptance length of 4.27 of 5. Speculation proposes several tokens per decode step and verifies them in one forward pass, so when the output repeats the prompt most proposed tokens are accepted at once, without changing what the model would have generated.

This variant was run end to end on GKE on the same single-L4 platform shape; the ModelDeployment below is the exact manifest from that run, and the numbers above are from the same run. Apply it instead of (or alongside) the deployment above:

model-deployment.yaml
# Qwen3-8B served on a single NVIDIA L4 by vLLM with n-gram (prompt-lookup)
# speculative decoding, validated end to end (the model layer is cloud-agnostic;
# the same manifest serves on EKS and GKE).
#
# n-gram speculation proposes the next tokens by matching a short suffix of what
# has been generated so far against earlier text in the prompt, then verifies the
# guess in one forward pass. It needs no draft model and no second set of weights,
# so it stays a single Standalone engine with no ModelCache. That is deliberate:
# Modelplane cannot yet stage a separate draft model on cache (modelplaneai/
# modelplane#281), so this is the speculative flavor that works today.
#
#   --speculative-config            method=ngram with num_speculative_tokens=5
#                                   proposes up to 5 tokens per step;
#                                   prompt_lookup_min/max=2..4 set the n-gram
#                                   suffix lengths matched against the prompt.
#                                   It pays off only when output repeats the
#                                   input - e.g. editing a pasted code block,
#                                   where most output tokens are copied verbatim.
#   --default-chat-template-kwargs  turns thinking off. Qwen3 thinks by default,
#                                   and a <think> block is novel text absent from
#                                   the prompt, so prompt-lookup cannot accelerate
#                                   it. Off, the output is mostly the copied code,
#                                   which is exactly what n-gram speeds up.
#   --max-model-len / --gpu-memory-utilization  L4 fit, not correctness. n-gram
#                                   adds only a small proposal buffer, no weights,
#                                   so the budget matches the plain Qwen3-8B recipe.
#
# No --port or --host: Modelplane's routing expects the engine on its default
# :8000 with a /health probe, and passes args through verbatim.
apiVersion: modelplane.ai/v1alpha1
kind: ModelDeployment
metadata:
  name: qwen3-8b-spec
  namespace: ml-team
spec:
  # One replica, matched to any compatible InferenceCluster by device capacity.
  replicas: 1
  template:
    spec:
      engines:
      - name: qwen3-8b-spec
        members:
    # A single self-contained vLLM pod. The container named "engine" is the
    # inference server; its image and args pass through verbatim.
        - role: Standalone
          nodeSelector:
            devices:
            - name: gpu
              count: 1
              selectors:
          # An 8B model needs most of an L4. >=20Gi selects the L4 (which reports
          # ~23Gi) without over-constraining. DRA evaluates this CEL against the
          # InferenceClass device, then against the GPU's ResourceSlice on bind.
              - cel: |
                  device.capacity["gpu.nvidia.com"].memory.compareTo(quantity("20Gi")) >= 0
          template:
            spec:
              containers:
              - name: engine
                image: vllm/vllm-openai:v0.23.0
                args:
                - "--model=Qwen/Qwen3-8B"
            # The id clients pass as "model" in OpenAI requests.
                - "--served-model-name=qwen3-8b-spec"
            # Cap the context so the KV cache fits beside the weights on the L4.
                - "--max-model-len=16384"
                - "--gpu-memory-utilization=0.92"
            # Enable n-gram speculative decoding (no draft model, no cache).
                - "--speculative-config={\"method\": \"ngram\", \"num_speculative_tokens\": 5, \"prompt_lookup_max\": 4, \"prompt_lookup_min\": 2}"
            # Thinking off, so output copies the prompt and prompt-lookup pays off.
                - "--default-chat-template-kwargs={\"enable_thinking\": false}"
model-service.yaml
# Exposes the qwen3-8b-spec deployment's endpoints as a single OpenAI-compatible
# URL. Modelplane labels each composed ModelEndpoint with the deployment name, so
# this selector reaches every replica. Read the public address from
# status.address:
#   kubectl get ms qwen3-8b-spec -n ml-team -o jsonpath='{.status.address}'
apiVersion: modelplane.ai/v1alpha1
kind: ModelService
metadata:
  name: qwen3-8b-spec
  namespace: ml-team
spec:
  endpoints:
  - selector:
      matchLabels:
        modelplane.ai/deployment: qwen3-8b-spec

Speculation is active when the engine logs its SpeculativeConfig at startup (method='ngram'). The call below pastes a code block and asks for a small edit, the copy-heavy case n-gram accelerates, so most output tokens are matched straight from the prompt:

bash
ADDR=$(kubectl get ms qwen3-8b-spec -n ml-team -o jsonpath='{.status.address}')
curl -s "$ADDR/v1/chat/completions" -H 'Content-Type: application/json' -d '{
  "model": "qwen3-8b-spec",
  "messages": [{"role":"user","content":"Return this Python function unchanged except rename the variable `total` to `subtotal`. Output only the code.\n\ndef cart(items):\n    total = 0\n    for item in items:\n        total += item.price\n    return total"}],
  "max_tokens": 200, "temperature": 0 }'

With the engine running, its logs report how many proposed tokens it accepts:

bash
kubectl logs -n ml-team -l modelplane.ai/deployment=qwen3-8b-spec \
  | grep "SpecDecoding metrics"