agent-platform release v4.48.1
Fixed
- (model-serving) Gemma-4-31b serves 32k and qwen3-6-35b-a3b 16k on the 48 GB card — the 64k context needed 12 GiB of KV cache next to 31.7 GiB of weights and vLLM refused to start (#591) in #607 by @teemow
Full Changelog: https://github.com/giantswarm/agent-platform/compare/v4.48.0...v4.48.1