Modelplane Modelplane docs

All recipes

Every validated recipe in one table: model, size, architecture, precision, and the verified hardware. Select a row for the full recipe.

ModelSizeArchPrecisionVerified onNotes
Qwen3-8B qwen8BDenseBF16EKS L4An 8.2B dense chat model on a single NVIDIA L4.
Qwen3-Coder-480B qwen480B A35BMoEBF16 / FP8EKS H200A 480B code MoE, multi-node BF16 over EFA or single-node FP8 on SGLang.
Kimi-K2 moonshotai1T A32BMoEINT4EKS H200A 1T MoE served prefill/decode disaggregated across two H200 nodes.
Llama-3.1-8B meta-llama8BDenseBF16EKSGKE L4An 8B dense chat model on a single NVIDIA L4.
GLM-4.5-Air zai-org106B A12BMoEGGUF IQ4_XSGKE A100A 106B MoE served from a GGUF checkpoint via llama.cpp on a single A100.
Nemotron-3.5-Lightning nvidia30B A3BMoENVFP4Nebius H100An open 30B MoE with 3B active parameters served NVFP4 on a single H100 on Nebius.