Pick a model and a workload. The estimate covers weights, KV cache for the context and concurrency you need, and runtime overhead, then ranks every Runpod GPU that fits by live price, with the Serverless versus always-on Pod break-even.
Show the math and edit the architecture
GPU options that fit
| GPU | Count | Fits up to | Pod $/hr | Pod $/month | Serverless $/day | Break-even busy | Stock |
|---|
Pod month = 730 hours. GPU count respects each GPU's per-machine limit for the chosen cloud. Serverless = flex worker price per GPU-hour from runpod.io/pricing () times worker-busy hours, so idle time costs nothing. Break-even busy = share of the day a worker must be busy before an always-on Pod is cheaper. Active-worker discounts are quoted by sales and not modelled. Click a row to load it into the PoC kit.
Where an APAC customer's workload can actually run today. Live datacenter list and per-GPU stock from Runpod's API, ranked by distance from the customer's city, with network-volume and S3 API support, which decide where data can live.
Datacenters by distance
| Datacenter | Location | Distance | RTT floor | Storage flag | S3 API | GPU stock |
|---|
RTT floor = great-circle distance at fibre speed (about 200 km per ms) with a 1.5x route factor. It is a floor for planning; measure from the customer's network before committing. City for a region code is inferred where Runpod publishes only the country. Storage flag = the API's storageSupport field; its exact meaning is not documented (some S3 API datacenters show it as false), so confirm volume support in the console.
Paste a Serverless endpoint ID and API key, or load a sample snapshot. The doctor reads /health, explains what the worker and job counts mean, and runs a timed test request. Calls go from your browser to api.runpod.ai directly.
Live endpoint
Full check reads /health (browser to api.runpod.ai), then the endpoint's v2 configuration, workers, network volumes and Serverless GPU availability (read-only, through this site's worker), the model's Hugging Face metadata, recent status incidents, and the logs of the least healthy worker. Use a read-only key and revoke it after the session.
Or load a sample snapshot
Samples use the documented /health shape and worker states (initializing, idle, running, throttled, unhealthy).
No reading yet.
Paste worker or job logs from a ticket. Each line is matched against known failure signatures; every hit gives the cause, the fix, who owns it, and when to escalate.
REST API v1 is retired on 2026-11-15 and GraphQL follows in early 2027. Paste a customer's integration code; the check flags every v1 call and field, rewrites what is mechanical, and lists what needs a human.
Customer code
Rewritten (mechanical changes only)
A proof-of-concept kit for the configuration picked in tab 1: the v2 API call that creates the endpoint, the client code the customer runs, a load test, and the success criteria agreed before the PoC starts.