draforge

DRAForge

DRAForge is a provider-neutral Kubernetes Dynamic Resource Allocation (DRA) observability, simulation, and diagnostics platform for GPU and accelerator workloads.

It helps platform engineers model virtual device pools, inspect ResourceClaim and ResourceSlice state, and explain why dynamic resource allocations succeed or fail—without requiring physical accelerator hardware or a specific cloud provider.

Latest release · GitHub repository · Discussions · Security policy · Support

Start here

Goal Guide
Run the verified local kind stack Installation guide
Install or operate DRAForge on Kubernetes Installation guide and operations guide
Understand components and data flow Architecture
Check supported Kubernetes DRA behavior Kubernetes DRA compatibility
Diagnose pending claims or runtime failures Troubleshooting guide
Explore the dashboard, CLI, and API Dashboard guide and public API/CLI surface

What DRAForge includes

Documentation map

Users and operators

Contributors and maintainers

Project scope

DRAForge components use Kubernetes APIs and Helm contracts rather than provider-specific services. They can run on compatible managed Kubernetes offerings, self-managed clusters, on-premises environments, and local test clusters. Provider-specific infrastructure and registry assets in this repository are optional showcases.

Keywords

Kubernetes DRA, Dynamic Resource Allocation, ResourceClaim, ResourceSlice, device plugin, GPU simulator, accelerator simulator, observability, diagnostics, Kubernetes troubleshooting, managed Kubernetes, on-premises Kubernetes, kind, Go, React, TypeScript, Helm, Terraform.