warply-ai/warply

Open-source control plane for AI endpoints.

Control your token factory.

Launch, inspect, and optimize OpenAI-compatible model serving on your GPUs. Warply turns a Python spec into runnable inference infrastructure across providers and backends.

GitHub

Launch

Create OpenAI-compatible endpoints from Python — local mock, SkyPilot provisioning, and SGLang backend adapters.

Observe

Inspect deployment status, compiled plans, and exported YAML today — with stats, traces, and pool metrics on the roadmap.

Optimize

Tune GPU placement, prefill/decode pools, and speculative decoding configs — compiled into a portable DeploymentPlan.

See how Warply routes a request across GPU pools

There are plenty of AI researchers and SWEs. There are nowhere near enough inference engineers.

Warply is a Python control plane for AI endpoints — launch OpenAI-compatible serving on your GPUs and clouds without Kubernetes CRDs, YAML, or per-cloud glue.

Portable control plane — compared.

Frameworks solve pieces of the stack. Warply unifies a Python-first control plane — cloud portability, observability hooks, and serving optimizations in one SDK you can move across providers.

Your token factory.
Your cloud. Your Python.

Warp-level control — launch, observe, and optimize from one SDK.

Install Warply and launch OpenAI-compatible model serving in a few lines — portable across providers when you need to move.