Distributed LLM Inference on Kubernetes: Configuring KServe and llm-d
Open Weight Models (Qwen, DeepSeek, Kimi) need a place to run that gives the ability to have shareable GPUs, a first-class orchestration/scheduler, and traffic routing capabilities that engineers are comfortable with.