Independent prototype built for a job application. Not affiliated with, endorsed by, or operated by Runpod. Prices, stock, incidents and the v2 schema are read live from Runpod's public sources. In the Endpoint doctor, health and test requests go from your browser straight to api.runpod.ai; configuration reads go through this site's worker as read-only GETs with secrets redacted, and the key is never logged or stored.

APAC FDE Workbench

The tools a forward deployed engineer uses on a Runpod pre-sales call and an escalation: size a model, price it live, pick an APAC-reachable datacenter, debug a Serverless endpoint, read its logs, move a customer off REST v1 before 2026-11-15, and hand over a PoC kit.

Pick a model and a workload. The estimate covers weights, KV cache for the context and concurrency you need, and runtime overhead, then ranks every Runpod GPU that fits by live price, with the Serverless versus always-on Pod break-even.

Show the math and edit the architecture

GPU options that fit

GPUCountFits up toPod $/hrPod $/monthServerless $/dayBreak-even busyStock

Pod month = 730 hours. GPU count respects each GPU's per-machine limit for the chosen cloud. Serverless = flex worker price per GPU-hour from runpod.io/pricing () times worker-busy hours, so idle time costs nothing. Break-even busy = share of the day a worker must be busy before an always-on Pod is cheaper. Active-worker discounts are quoted by sales and not modelled. Click a row to load it into the PoC kit.

Where an APAC customer's workload can actually run today. Live datacenter list and per-GPU stock from Runpod's API, ranked by distance from the customer's city, with network-volume and S3 API support, which decide where data can live.

Datacenters by distance

DatacenterLocationDistanceRTT floorStorage flagS3 APIGPU stock

RTT floor = great-circle distance at fibre speed (about 200 km per ms) with a 1.5x route factor. It is a floor for planning; measure from the customer's network before committing. City for a region code is inferred where Runpod publishes only the country. Storage flag = the API's storageSupport field; its exact meaning is not documented (some S3 API datacenters show it as false), so confirm volume support in the console.

Paste a Serverless endpoint ID and API key, or load a sample snapshot. The doctor reads /health, explains what the worker and job counts mean, and runs a timed test request. Calls go from your browser to api.runpod.ai directly.

Live endpoint

Full check reads /health (browser to api.runpod.ai), then the endpoint's v2 configuration, workers, network volumes and Serverless GPU availability (read-only, through this site's worker), the model's Hugging Face metadata, recent status incidents, and the logs of the least healthy worker. Use a read-only key and revoke it after the session.

Or load a sample snapshot

Samples use the documented /health shape and worker states (initializing, idle, running, throttled, unhealthy).

No reading yet.

Paste worker or job logs from a ticket. Each line is matched against known failure signatures; every hit gives the cause, the fix, who owns it, and when to escalate.

REST API v1 is retired on 2026-11-15 and GraphQL follows in early 2027. Paste a customer's integration code; the check flags every v1 call and field, rewrites what is mechanical, and lists what needs a human.

Customer code

Rewritten (mechanical changes only)


    

A proof-of-concept kit for the configuration picked in tab 1: the v2 API call that creates the endpoint, the client code the customer runs, a load test, and the success criteria agreed before the PoC starts.

1. Create the endpoint (REST v2)

2. Call it (OpenAI-compatible)

3. Load test (TTFT and throughput)

4. Success criteria, agreed on day 0