-
Serving many models on one GPU: Triton, model repositories, and hot-swapping
A GPU is not a Kubernetes pod. Treating every model as if it deserves an entire accelerator can leave expensive hardware mostly idle, especially when you serve small classifiers, regressors, ranking m