agent-platform release v4.48.2
Fixed
- (model-serving) Gemma-4-31b serves 8k on the 48 GB card — vLLM’s measured ceiling next to 31.7 GiB of weights; 32k returns with an fp8 KV cache once proven on the L40S (#591) in #612 by @teemow
Full Changelog: https://github.com/giantswarm/agent-platform/compare/v4.48.1...v4.48.2