Start typing to search across invoices, services, domains, tickets, and more...
If you already own GPU servers, the cheapest long-term way to serve AI inference in Asia or the US is usually to colocate them in a data center close to your users rather than renting cloud GPUs by the hour. IMIDC offers colocation from 1U to a full rack in Tokyo, Japan; Hong Kong; Singapore and Los Angeles, USA, with remote hands, 10Gbps uplinks, IP blocks and BGP with your own ASN — note that IMIDC does not rent GPUs; this is bring-your-own hardware.
Colocation wins when your GPUs run most of the day for a year or more; cloud rental wins for short bursts and experiments.
| Factor | Colocating your own GPUs (IMIDC) | Cloud GPU rental |
|---|---|---|
| Cost structure | Hardware capex + monthly space, power and bandwidth | Hourly or reserved rate, no capex |
| Best utilisation profile | Steady 24/7 inference | Spiky, short-term or training bursts |
| Hardware choice | Any GPU, CPU, RAM and NVMe you buy | Provider's catalogue and availability |
| Data location | Your hardware in a known city | Provider region, sometimes shared tenancy |
| Network identity | Your own IP block and ASN via BGP | Provider IPs |
| Scaling speed | Weeks (procure, ship, install) | Minutes, if capacity exists |
| Failure handling | You keep spares; remote hands swap parts | Provider replaces instance |
A common pattern is hybrid: baseline inference on colocated servers, overflow and experiments on rented capacity.
Power, not rack units, is the real constraint for GPU colocation, so confirm the kW you need before buying space.
| Typical chassis | Example GPU load | Rough power draw at load | Colocation option |
|---|---|---|---|
| 1U server | 1-2 low-profile inference GPUs | about 0.5-1 kW | 1U-2U space |
| 2U server | 2-4 PCIe GPUs | about 1.5-3 kW | 2U-4U or partial rack |
| 4U server | 8 PCIe or SXM GPUs | about 5-10+ kW | Full rack, often one or two servers per rack |
| Several 4U servers | Cluster for larger models | Tens of kW | Multiple racks; plan with sales |
These are typical ranges, not IMIDC specifications — read your vendor's nameplate and measure under real load. Practical points:
# measure real draw under load before you size the rack
nvidia-smi --query-gpu=index,name,power.draw,power.limit,temperature.gpu --format=csv -l 5
# cap each GPU so the whole chassis fits the contracted power budget
sudo nvidia-smi -pm 1 # persistence mode
sudo nvidia-smi -i 0 -pl 300 # example: limit GPU 0 to 300 W
# whole-server reading from the BMC (IPMI)
ipmitool -I lanplus -H 198.51.100.20 -U admin -P '***' dcmi power reading
Inference traffic is small per request but latency-sensitive, so prioritise clean routing and your own address space over raw bandwidth.
# example: serve an open-weight model with vLLM, API on localhost only
docker run -d --name vllm --gpus all --restart unless-stopped \
-p 127.0.0.1:8000:8000 --ipc=host \
-v /data/models:/models \
vllm/vllm-openai:latest \
--model /models/your-model --tensor-parallel-size 4
curl -s http://127.0.0.1:8000/v1/models
Pick the city nearest to the majority of requests, then test from where your users actually are.
| IMIDC location | Typical audience | Notes |
|---|---|---|
| Tokyo, Japan | Japan, Korea, wider North Asia | Native Japanese IPs; strong regional peering |
| Hong Kong | Greater China, Southeast Asia | CN2 GIA to mainland China; check export-control rules for advanced GPUs |
| Singapore | Southeast Asia, India, Oceania | Configured via sales |
| Los Angeles, USA | North America, trans-Pacific | Unicom 9929/4837 and CN2 routes to China |
# from a user-side test host, compare candidate locations
mtr -rwc 100 203.0.113.10 # Tokyo test IP (example)
mtr -rwc 100 203.0.113.20 # Singapore test IP (example)
Advanced AI chips are subject to export-control rules, and you, as owner and shipper, are responsible for compliance.
This is a brief overview, not legal advice. Consult a trade-compliance specialist.
Most delays come from paperwork and missing parts, not from racking.
After install, remote hands can handle GPU or disk swaps from your spares, reseats and power cycles via ticket.
Match the colocation size to measured power, then add network services.
High-density power, Singapore colocation and multi-rack projects are custom configurations quoted by IMIDC sales.
IMIDC offers colocation from 1U to a full rack in Tokyo, Hong Kong, Singapore and Los Angeles, with remote hands and 10Gbps uplinks. Confirm your power requirement with sales before shipping hardware.
No. IMIDC provides colocation space, power, network and IP resources for hardware you own. You can pair colocated GPUs with IMIDC dedicated servers for CPU workloads.
Typically several kW and sometimes over 10 kW at full load, depending on the GPU model. Measure under real load, consider power caps, and plan a full rack with sales.
Yes. IMIDC supports BGP announcement with your own ASN, IPv4 leasing including full /24 blocks, and Anycast.
Planning a GPU deployment? Review IMIDC colocation and data centers, BGP and Anycast and IP resources, then contact sales with your server model and power figures, or open a ticket to schedule an install.