Application layer¶
Pulumi Go stack that deploys the Shaide application into Kubernetes namespace app-shaide.
This stack owns the application layer (Shaide server, control panel UI, RustFS object storage,
and Qdrant vector database) and wires everything together with shared configuration and
secrets managed through Pulumi.
Repository Layout¶
app_shaide/ itself only contains main.go (a thin wrapper that calls
shaide.DeployAppShaide(ctx)) and deployments/. The actual implementation lives in the
shared pkg module, at pkg/iac/shaide/:
pkg/iac/shaide/
├── shaide.go # DeployAppShaide — orchestration entry point
└── internal/
├── config/config.go # Config loading; typed view of Pulumi stack config
├── runtime/context.go # DeploymentContext — shared labels + dependency options
├── platform/
│ ├── k8s-serviceaccount.go # shaide-server ServiceAccount (+ workload-identity annotations)
│ ├── k8s-rbac.go # ClusterRole/ClusterRoleBinding — cluster-wide pod watch
│ ├── configmap.go # shaide-config ConfigMap + shaide-secrets Secret
│ └── secret.go # ghcr-creds pull secret (ghcr.io or Harbor)
├── components/
│ ├── shaide/deploy.go # shaide-server StatefulSet + Service (+ HTTPRoute)
│ ├── controlpanel/deploy.go # control-panel Deployment + Service
│ ├── webapp/deploy.go # webapp Deployment + Service
│ ├── rustfs/deploy.go # rustfs StatefulSet + Service
│ └── qdrant/deploy.go # qdrant StatefulSet + Service
└── cloudprovider/ # Provider interface + per-cloud implementations
├── provider.go, gcp.go, aws.go, azure.go, on-prem.go, generic.go
What Each Component Does¶
main.go¶
Compiles the stack binary and delegates to shaide.DeployAppShaide(ctx).
pkg/iac/shaide/shaide.go¶
The stack orchestrator and dependency coordinator:
- Loads Pulumi config via appconfig.Load(ctx).
- Creates the Kubernetes provider used by all resources in this stack.
- Creates the namespace first, so all subsequent resources are scoped correctly.
- Creates shared prerequisites: GHCR/Harbor pull secret, shaide-config ConfigMap,
shaide-secrets Secret, ServiceAccount, and cluster-wide RBAC.
- Selects a cloudprovider.Provider (see below) and calls ProvisionStorage before any
StatefulSet is created.
- Calls per-component Deploy functions in a stable order: shaide, controlpanel,
webapp, rustfs, qdrant.
- Threads shared dependencies through a runtime.DeploymentContext so resources only
create after their inputs are ready.
pkg/iac/shaide/internal/platform/¶
k8s-serviceaccount.go: Creates the ServiceAccount used byshaide-server(shaideServiceAccountName). Attaches whatever annotations are set inserviceAccountAnnotations— a generic map, so the same code handles GKE Workload Identity, AKS Workload Identity, EKS IRSA, or nothing at all (on-prem/generic).k8s-rbac.go: Creates a cluster-scopedClusterRole/ClusterRoleBindinggrantingshaide-server's ServiceAccountget/list/watchonpodscluster-wide — this is what lets Shaide observe pod state across every namespace it needs to (its own,app_serving's per-model namespaces,app_mcp's namespace, etc.), not just its own namespace.configmap.go: Creates the sharedshaide-configConfigMap (non-secret runtime settings, service discovery) andshaide-secretsSecret (adminAuthKey,s3Password).secret.go: Createsghcr-creds(kubernetes.io/dockerconfigjson). Authenticates againstghcr.iousingghcrUser/ghcrToken— or, whenharborHostnameis set, against that internal Harbor registry instead, using the same two config keys.
pkg/iac/shaide/internal/components/¶
shaide/deploy.go: Deploysshaide-serveras a StatefulSet with a PVC mounted at/root/.configfor SQLite persistence, and exposes a Service. The Service isClusterIP(paired with an HTTPRoute to a shared Gateway) wheneverinfraStackReforgatewayHostnameis set; otherwise it's aLoadBalancerwith annotations fromlbAnnotations. Delegates cloud-specific post-deploy resources to the activecloudprovider.Provider.controlpanel/deploy.go: Deploys the control panel as a Deployment (single replica, no persistence) with aClusterIPService on port3000.webapp/deploy.go: Deploys the end-user facing web application as a Deployment (single replica, no persistence) with aClusterIPService on port8787. Same shape as the control panel — internal-only, points atshaide-serverviaSHAIDE_SERVER_FQDN/SHAIDE_SERVER_PORTenv vars.rustfs/deploy.go: Deploysrustfsas a StatefulSet with a PVC for/data, anemptyDirfor/logs, and aClusterIPService on port9000(plus9001whenrustfsConsoleEnabledis set). The container runs as UID/GID10001and is configured via values from the shared ConfigMap and Secret.qdrant/deploy.go: Deploysqdrantas a StatefulSet with a PVC for/qdrant/storageand aClusterIPService exposing REST6333and gRPC6334.
pkg/iac/shaide/internal/cloudprovider/¶
A Provider interface (provider.go) with one implementation per target, selected by
the informational cloudProvider config value:
cloudProvider |
ProvisionStorage |
PostDeployService |
|---|---|---|
gcp |
no-op (GKE pd.csi.storage.gke.io dynamic provisioner) |
Creates a GKE HealthCheckPolicy targeting /v1/health |
azure |
no-op (AKS disk.csi.azure.com dynamic provisioner) |
no-op (Workload Identity pod label is applied directly in shaide/deploy.go) |
aws |
no-op (EBS CSI dynamic provisioner) | no-op |
on-prem |
Creates one static, pre-bound hostPath PersistentVolume per stateful component (shaide-server, rustfs, qdrant), pinned to pvNodeHostname — only when storageClassName is hostpath |
no-op (MetalLB handles LB via the lbAnnotations Service annotation) |
| anything else | no-op | no-op |
Any unrecognized cloudProvider value falls back to the generic no-op provider — useful
for local/dev clusters that already have a working default StorageClass and don't need a
LoadBalancer at all.
Component Ports¶
| Component | Service Name | Ports | Scope |
|---|---|---|---|
| Shaide server | shaide-server |
80 -> 8080 |
External (LoadBalancer) or internal (Gateway/HTTPRoute) |
| Control panel | control-panel |
3000 |
Internal only |
| Web app | webapp |
8787 |
Internal only |
| RustFS | rustfs |
9000 (9001 when app_shaide:rustfsConsoleEnabled=true) |
Internal only |
| Qdrant | qdrant |
6333, 6334 |
Internal only |
emptyDir Permissions Note (RustFS)¶
emptyDir does not support direct permission or ownership settings in pod spec. The fix is
an initContainer (defined in pkg/iac/shaide/internal/components/rustfs/deploy.go) that prepares the filesystem
before the main container starts:
- Runs as root.
- Sets /data and /logs to mode 0755.
- Sets ownership to 10001:10001.
- Starts the main rustfs container after permissions are correct.
This matches RustFS expectations (0o755) and avoids runtime permission errors.
Configuration Source of Truth¶
All parameters are defined in the active stack's deployments/Pulumi.<stack>.yaml. Each stack config is the authoritative list of settings for that deployment target, including:
- Namespace and platform behavior (app_shaide:namespace, app_shaide:cloudProvider,
app_shaide:infraStackRef, app_shaide:nodeSelector, app_shaide:shaideServiceAccountName).
- Container images (shaideServerImage, controlPanelImage, webappImage, rustfsImage, qdrantImage).
- Runtime config (shaideServerS3Fqdn, shaideServerS3Port, databaseUrl, vectorDBUrl).
- Trial deployment flag (trial, defaults to FALSE; injected into shaide-server as the
TRIAL env var — only the trial stack sets it to TRUE).
- RustFS console exposure (rustfsConsoleEnabled).
- MCP integration (mcpNamespace, optional — see MCP Integration).
- Secrets (ghcrToken, adminAuthKey, s3Password).
The nodeSelector value is rendered as a soft (preferred, not required) nodeAffinity for a
nodegroup label on the target node pool (e.g. shaide-nodepool) — components prefer a
matching node but can still schedule elsewhere if none is available. Each component also gets
a soft podAntiAffinity spreading its own replicas across nodes.
Prereqs¶
- Kubernetes context points to the target cluster.
- Pulumi stack is selected for this project.
- Required secrets are set (
ghcrToken,adminAuthKey,s3Password).
Workload Identity¶
app_shaide:serviceAccountAnnotations is a generic annotation map applied to the
shaide-server ServiceAccount — the same mechanism works for any cloud's workload
identity binding:
- GKE Workload Identity:
iam.gke.io/gcp-service-account: <gsa-email> - AKS Workload Identity:
azure.workload.identity/client-id: <managed-identity-client-id>(Azure additionally requires theazure.workload.identity/use: "true"pod label, which is applied automatically whenevercloudProvider: azure.) - EKS IRSA:
eks.amazonaws.com/role-arn: <role-arn> - On-prem/generic: omit
serviceAccountAnnotationsentirely.
app_shaide:shaideServiceAccountName (defaults to shaide-server) names the
ServiceAccount these annotations are applied to.
Required IAM on the mapped cloud identity (for Vertex AI access): roles/aiplatform.user
or the equivalent role on the target cloud.
MCP Integration (optional)¶
app_shaide:mcpNamespace points shaide-server at the namespace where the app_mcp
stack is deployed (e.g. mcp-gateway). It is optional:
- If set, MCP_NAMESPACE is injected into shaide-config and shaide-server can
discover/watch MCP pods in that namespace.
- If left unset, MCP_NAMESPACE is omitted from shaide-config entirely and
shaide-server runs without MCP support — no app_mcp deployment is required.
Security Notes¶
- Sensitive values live in the
shaide-secretsKubernetes Secret created by Pulumi. - The GHCR token must have
read:packagesto pull the private Shaide image. - Avoid committing plaintext secrets into stack config; use
pulumi config set --secret.
Resource Ownership¶
This stack owns only the app-shaide application layer resources. Cluster-level routing,
Gateways, and cloud infra are expected to be managed by the infra stacks referenced via
infraStackRef.
Data Persistence¶
Persistent storage is provided by PVCs for Shaide SQLite (/root/.config), RustFS
(/data), and Qdrant (/qdrant/storage). Deleting PVCs will permanently remove stored data.
Gateway Mode¶
Routing mode is cloud-agnostic — it depends only on whether a Gateway hostname is
available, not on cloudProvider. When either infraStackRef (a StackReference to an
infra stack that exports gatewayHostname) or gatewayHostname (set directly) is
non-empty, the Shaide Service becomes ClusterIP and an HTTPRoute is created to attach
it to the shared Gateway. When neither is set, Shaide is exposed directly via a
LoadBalancer Service with annotations from lbAnnotations. The two are mutually
exclusive in practice — set one or the other, not both.
Config Changes¶
Update the active stack config (deployments/Pulumi.<stack>.yaml), then apply changes with
pulumi up to reconcile the cluster.
Mirror Images (on-prem)¶
On-prem deployments cannot pull from ghcr.io directly. Images must be mirrored into
Harbor from the provisioner laptop before running pulumi up.
1 — Authenticate with GHCR¶
The Shaide images are in a private GitHub Container Registry package. Authentication
requires a GitHub Personal Access Token (PAT) with read:packages scope.
Option A — use the gh CLI (recommended):
gh auth login # follow the prompts; select HTTPS + browser
gh auth token # prints the token — used below
Then authenticate skopeo:
skopeo login ghcr.io \
--username "$(gh api user --jq .login)" \
--password "$(gh auth token)"
Option B — use a PAT directly:
skopeo login ghcr.io \
--username <your-github-username> \
--password <your-PAT>
Credentials are cached in ~/.config/containers/auth.json for subsequent skopeo calls.
2 — Download images¶
skopeo copy docker://ghcr.io/axem-solutions/shaide_server:v0.7.0 \
oci-archive:infra/on-prem/ansible/artifacts/images/shaide_server-v0.7.0.tar
skopeo copy docker://ghcr.io/axem-solutions/control_panel:v0.3.0 \
oci-archive:infra/on-prem/ansible/artifacts/images/control_panel-v0.3.0.tar
3 — Upload to Harbor¶
cd infra/on-prem/ansible
ansible-playbook -i inventory-dev harbor_upload.yml
Deploy¶
cd app_shaide
pulumi up
Quick Checks¶
# Namespace and workloads
kubectl get ns app-shaide
kubectl get pods -n app-shaide -o wide
# Services and endpoints
kubectl get svc -n app-shaide -o wide
kubectl get endpoints -n app-shaide
# ServiceAccount + WI mapping (GKE)
kubectl get pod -n app-shaide shaide-server-0 -o jsonpath='{.spec.serviceAccountName}{"\n"}'
kubectl get sa -n app-shaide shaide-server -o yaml | rg "iam.gke.io/gcp-service-account"
kubectl exec -n app-shaide shaide-server-0 -- sh -lc \
'curl -s -H "Metadata-Flavor: Google" http://metadata.google.internal/computeMetadata/v1/instance/service-accounts/default/email; echo'
# Shaide health (inside cluster)
# Replace service names if you customized them in Pulumi config.
kubectl run -n app-shaide shaide-health --rm -i --restart=Never --image=curlimages/curl:8.5.0 -- \
curl -sS http://shaide-server/v1/health
# RustFS permissions and logs
kubectl exec -n app-shaide rustfs-0 -- ls -ld /data /logs
kubectl logs -n app-shaide rustfs-0
Troubleshooting¶
ImagePullBackOfffor shaide-server:- Confirm
ghcr-credsexists inapp-shaideand contains validghcrUser/ghcrToken. -
Re-set the token with
pulumi config set --secret ghcrToken <token>andpulumi up. -
RustFS fails with permission errors:
- Ensure the
fix-permissionsinitContainer ran and set/dataand/logsto0755with10001:10001. -
Check initContainer logs:
kubectl logs -n app-shaide rustfs-0 -c fix-permissions. -
Shaide server not reachable:
- Verify
shaide-serverService type and endpoints:kubectl get svc,endpoints -n app-shaide. -
If on GCP with
infraStackRef, confirm HTTPRoute exists:kubectl get httproute -n app-shaide. -
Vertex returns
PERMISSION_DENIED: - Verify
shaide-serveruses the expected KSA and that KSA is annotated withiam.gke.io/gcp-service-account. -
Ensure the mapped GSA has
roles/aiplatform.user. -
Qdrant not responding:
- Check pod status and PVC binding:
kubectl get pods,pvc -n app-shaide. - Confirm ports are open in the Service:
kubectl describe svc -n app-shaide qdrant.