This guide covers day-two operations for managing a DRAForge deployment, including upgrades, observability (metrics and logs), and cleanup procedures.
Upgrading DRAForge primarily involves updating the Helm release.
Update the Helm Repository/Chart:
Ensure you have the latest chart definitions. If you are using a remote repository, run helm repo update. If using the source directory, ensure you have pulled the latest changes.
Run the Upgrade:
helm upgrade draforge deploy/helm/draforge \
--namespace draforge-system
v0.3 exposure change: External Gateway/Ingress resources are now disabled by default. Before upgrading an intentionally public deployment, prepare a secure exposure values file with TLS, an authentication proxy Service, and restrictive HTTPS CORS. The local-demo profile is explicitly unauthenticated and non-production.
kubectl get pods -n draforge-system -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[0].image}{"\n"}{end}'
Both the API server and simulation controller expose separate liveness and readiness endpoints:
/healthz is process-only liveness and remains HTTP 200 during transient Kubernetes API outages./readyz performs a read-only Kubernetes namespace list with a bounded request context. It returns a degraded HTTP 200 during the configured failure grace period and HTTP 503 if the dependency remains unavailable beyond that period.The public response uses stable reason codes and never includes raw Kubernetes client errors, credentials, API tokens, kubeconfig paths, or cluster URLs.
The Helm chart configures both workloads with these defaults:
server:
lifecycle:
readinessTimeout: 2s
readinessGracePeriod: 15s
shutdownTimeout: 5s
controller:
lifecycle:
readinessTimeout: 2s
readinessGracePeriod: 15s
shutdownTimeout: 5s
Set readinessGracePeriod: 0 to fail readiness immediately. Timeout and shutdown values must be positive Go-style durations using ms, s, m, or h. Keep shutdownTimeout below the Pod termination grace period (Kubernetes defaults it to 30 seconds) so the process can complete its graceful shutdown before kubelet sends SIGKILL.
draforge serve listens for SIGINT and SIGTERM. On shutdown it stops accepting new connections, cancels active request contexts and SSE streams, waits up to shutdownTimeout, and force-closes remaining HTTP connections if the graceful deadline expires. The controller uses the same signal-aware readiness and shutdown settings for its runtime server.
For direct binary use, the corresponding flags are:
--readiness-timeout
--readiness-grace-period
--shutdown-timeout
The controller uses a namespaced Kubernetes Lease by default. Every replica keeps
its /healthz, /readyz, and /metrics endpoints available, but only the current
lease holder starts reconciliation and simulated-allocation loops. The controller
Pod name is used as the lease identity through the Downward API.
controller:
replicaCount: 2
leaderElection:
enabled: true
leaseDuration: 15s
renewDeadline: 10s
retryPeriod: 2s
The chart rejects replicaCount values above one when leader election is disabled.
The corresponding binary flags are --leader-elect,
--leader-election-lease-name, --leader-election-lease-namespace,
--leader-election-identity, --leader-election-lease-duration,
--leader-election-renew-deadline, and --leader-election-retry-period.
draforge_controller_leader is 1 only on the active replica and 0 on standby
replicas. A process that loses leadership stops its active loops and exits so the
Deployment can restart it as a fresh candidate.
Leader leases are not released early during shutdown. This prevents an old leader
and a replacement from overlapping while in-flight work drains. During an upgrade,
rollback, abrupt Pod loss, or node disruption, the replacement may therefore wait
up to leaseDuration before becoming active. Keep the lease duration short enough
for the required recovery objective, while preserving leaseDuration >
renewDeadline > retryPeriod * 1.2. Roll back by restoring the previous chart/image
version; the same Lease name is retained, so only one version can be active at a
time.
The install-level compatibility gate runs two replicas, confirms the Lease holder
matches the only replica reporting draforge_controller_leader 1, deletes that
active Pod, and requires a different leader to reconcile a concurrent pool-update
burst and a new ResourceClaim without changing the original allocation.
The active controller uses Kubernetes shared informers instead of fixed polling intervals. Add, update, and delete events are coalesced into two typed workqueues:
Each queue has one worker, so pool passes are serialized with other pool passes
and allocation passes are serialized with other allocation passes; the two
pipelines can progress independently. A top-level synchronization error returned
by a worker increments draforge_controller_reconcile_errors_total. Conflict,
AlreadyExists, timeout, throttling, service-unavailable, internal, and unknown
errors are retried with exponential backoff starting at 100 milliseconds and
capped at 30 seconds. Retry history is cleared after success; otherwise the item
becomes terminal after eight retries. Authorization, validation, bad-request,
not-found, unsupported-method, and request-size errors are terminal immediately.
draforge_controller_reconcile_retries_total counts scheduled retries and
draforge_controller_terminal_errors_total counts terminal decisions.
ResourceSlice get/create/update/delete operations, SimulatedDevicePool status
writes, orphan-cleanup list/deletes, and ResourceClaim allocation-status writes
return their Kubernetes API errors to this queue policy instead of logging and
continuing. A delete that reports NotFound remains an idempotent success. The
allocation counter advances only after the ResourceClaim status write succeeds.
The current workers re-read authoritative API state on each coalesced pass;
informer events are the trigger, not a scheduler-equivalent cached allocation
model.
SimulatedDevicePool is namespaced while ResourceSlice is cluster-scoped, so
Kubernetes does not permit a valid owner reference from a ResourceSlice to its
pool. DRAForge therefore uses an explicit ownership contract instead:
draforge.oaslananka/resourceslice-cleanup finalizer;draforge.oaslananka/sdp-name,
draforge.oaslananka/sdp-namespace, and draforge.oaslananka/sdp-uid labels;Deletion failures retain the finalizer and return the Kubernetes API error to the bounded queue policy. A missing ResourceSlice is an idempotent success. UID matching prevents a deleting pool from removing a same-name slice owned by a different pool incarnation. During upgrade, an ordinary reconciliation pass adopts deterministic legacy slices by adding the namespace and UID labels. While finalizing a legacy pool, the controller may also remove its exact spec-derived deterministic slice names when the older UID labels are absent.
The controller ClusterRole has update and patch access to the main
SimulatedDevicePool resource only for finalizer management; status writes remain
in the separate status-subresource rule. Rollback to a pre-finalizer controller
should be performed only after managed pools have finished reconciliation or
been deleted, because an older binary cannot remove the new cleanup finalizer.
DRAForge provides native support for monitoring and logging to assist operators in maintaining cluster health.
The API server exposes Prometheus metrics on its HTTP service. The controller exposes runtime health and metrics on the named runtime port, which defaults to TCP 8082 and serves /metrics.
The controller metrics Service is enabled by default, but the controller NetworkPolicy keeps all ingress closed until a monitoring peer is explicitly selected. This prevents unrelated namespaces from scraping the controller merely because the Service exists.
Key controller lifecycle metrics include:
draforge_controller_leader: 1 on the active Lease holder and 0 on standby replicas.draforge_controller_reconcile_errors_total: synchronization attempts that returned an error.draforge_controller_reconcile_retries_total: rate-limited retries scheduled by the queue policy.draforge_controller_terminal_errors_total: errors forgotten immediately or after the retry limit.draforge_controller_sync_attempts_total{pipeline=...}: pool or allocation synchronizations started.draforge_controller_sync_duration_seconds_total{pipeline=...}: cumulative synchronization time for each pipeline.draforge_controller_sync_in_flight{pipeline=...}: synchronization workers currently executing.draforge_controller_queue_depth{pipeline=...}: ready items waiting in each workqueue; rate-limited delayed retries are not included until ready.Use a namespace selector and, preferably, a pod selector:
controller:
metrics:
networkPolicy:
enabled: true
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: monitoring
podSelector:
matchLabels:
app.kubernetes.io/name: prometheus
Install or upgrade with the values file:
helm upgrade --install draforge deploy/helm/draforge \
--namespace draforge-system \
--create-namespace \
-f monitoring-values.yaml
The generated NetworkPolicy permits only matching pods in matching namespaces to reach the controller container metrics port. Verify namespace labels with:
kubectl get namespace monitoring --show-labels
Leaving podSelector empty allows every pod in the selected namespace, so production deployments should normally provide both selectors. Disabling controller.metrics.service.enabled removes the Service and keeps the controller ingress policy closed even if the metrics NetworkPolicy values remain configured.
A ServiceMonitor can be rendered when the Prometheus Operator CRDs are already installed:
controller:
metrics:
serviceMonitor:
enabled: true
namespace: monitoring
labels:
release: prometheus
interval: 30s
scrapeTimeout: 10s
The ServiceMonitor selects the DRAForge controller Service in the Helm release namespace and scrapes the named runtime port at /metrics. It is disabled by default and schema validation rejects enabling it while the metrics Service is disabled. Enabling the ServiceMonitor does not open NetworkPolicy ingress; configure the restricted monitoring peer shown above as part of the same values profile.
DRAForge components log operational details to standard output, making them compatible with standard Kubernetes log aggregators (e.g., Fluent Bit, Promtail).
kubectl logs -l app.kubernetes.io/name=draforge-controller -n draforge-system
kubectl logs -l app.kubernetes.io/name=draforge-api-server -n draforge-system
Tip: Adjust log verbosity by setting the appropriate environment variables or flags defined in the component configurations if deeper debugging is required.
When you need to remove DRAForge or clean up test scenarios, follow these steps to ensure all resources are properly deleted.
Before uninstalling the core components, it’s good practice to remove any applied scenarios (like SimulatedDevicePools and ResourceClaims) to allow for graceful termination.
kubectl delete -f examples/scenarios/basic-gpu.yaml
# Or, clean up specific pools
kubectl delete simulateddevicepools --all
To completely remove the DRAForge deployment from your cluster:
helm uninstall draforge -n draforge-system
Helm does not automatically remove CRDs when a release is uninstalled. To fully clean up the cluster, you must remove the CRDs manually. Warning: This will delete all custom resources of this type in the cluster.
kubectl delete -f deploy/crds/simulateddevicepool-crd.yaml
If you deployed the DOKS showcase using the provided Terraform modules, refer to the Cost Control Guide for instructions on tearing down billable infrastructure using task demo:down or terraform destroy.