Model management¶
Listing¶
curl https://<endpoint>/v1/models -H "Authorization: Bearer <key>"
kubectl -n shaide-serving get pods -l app.kubernetes.io/component=modelservice
Adding¶
- Publish weights to the internal registry - see Model registry.
- Add the model definition under the serving stack's model deployments.
- Reference it in the stack configuration.
- Apply the serving stack.
Convention-based discovery: a gaie-<slug> / ms-<slug> folder pair under
deployments/models/<category>/<model>/ is picked up automatically once referenced in
stack config.
Detail: Model deployment flow.
Swapping on a single GPU¶
Where GPU capacity allows only one large model, scale the outgoing model to zero before
scheduling its replacement - otherwise the new pod stays Pending on insufficient VRAM.
Removing¶
Remove the model from stack configuration and re-apply. Weights remain in the internal registry until deleted there.
Verifying¶
kubectl -n shaide-serving get pods -w
curl https://<endpoint>/v1/models -H "Authorization: Bearer <key>"
A new model appears in the API list once at least one replica is ready.