ESC

Start typing to search across invoices, services, domains, tickets, and more...

Search... Ctrl+K
Use Cases & Solutions

Colocate Your Own GPU Servers for AI Inference in Tokyo, Hong Kong, Singapore or Los Angeles: Power, Network and Shipping Guide

7 steps 20 min read 15 views 0
On this page

If you already own GPU servers, the cheapest long-term way to serve AI inference in Asia or the US is usually to colocate them in a data center close to your users rather than renting cloud GPUs by the hour. IMIDC offers colocation from 1U to a full rack in Tokyo, Japan; Hong Kong; Singapore and Los Angeles, USA, with remote hands, 10Gbps uplinks, IP blocks and BGP with your own ASN — note that IMIDC does not rent GPUs; this is bring-your-own hardware.

Key facts
  • Locations for GPU colocation: Tokyo, Japan; Hong Kong; Singapore; Los Angeles, USA (Singapore arranged via sales).
  • Space from 1U to a full rack; power per rack must be confirmed with IMIDC sales before you ship.
  • Remote hands for racking, cabling, GPU/disk swaps and power cycles.
  • 10Gbps uplinks, DDoS protection, IPv4 leasing (full /24 available) and BGP announcement with your own ASN.
  • IMIDC does not sell or rent GPUs; export-control compliance for your hardware is your responsibility.

When colocation beats renting cloud GPUs

Colocation wins when your GPUs run most of the day for a year or more; cloud rental wins for short bursts and experiments.

FactorColocating your own GPUs (IMIDC)Cloud GPU rental
Cost structureHardware capex + monthly space, power and bandwidthHourly or reserved rate, no capex
Best utilisation profileSteady 24/7 inferenceSpiky, short-term or training bursts
Hardware choiceAny GPU, CPU, RAM and NVMe you buyProvider's catalogue and availability
Data locationYour hardware in a known cityProvider region, sometimes shared tenancy
Network identityYour own IP block and ASN via BGPProvider IPs
Scaling speedWeeks (procure, ship, install)Minutes, if capacity exists
Failure handlingYou keep spares; remote hands swap partsProvider replaces instance

A common pattern is hybrid: baseline inference on colocated servers, overflow and experiments on rented capacity.

Power density and cooling: size this first

Power, not rack units, is the real constraint for GPU colocation, so confirm the kW you need before buying space.

Typical chassisExample GPU loadRough power draw at loadColocation option
1U server1-2 low-profile inference GPUsabout 0.5-1 kW1U-2U space
2U server2-4 PCIe GPUsabout 1.5-3 kW2U-4U or partial rack
4U server8 PCIe or SXM GPUsabout 5-10+ kWFull rack, often one or two servers per rack
Several 4U serversCluster for larger modelsTens of kWMultiple racks; plan with sales

These are typical ranges, not IMIDC specifications — read your vendor's nameplate and measure under real load. Practical points:

  • Measure and cap: GPUs can be power-limited with almost no inference throughput loss at moderate caps.
  • Airflow: GPU servers are designed for front-to-back airflow; tell IMIDC the chassis depth and airflow direction so it is placed correctly in hot/cold aisles.
  • Redundant PSUs: ask for A+B power feeds and make sure the server can run on one feed at your capped load.
  • Plugs: high-wattage PSUs often use C19/C20; ship matching power cords.
  • Liquid cooling: discuss with sales before buying direct-liquid-cooled systems; do not assume support.
# measure real draw under load before you size the rack
nvidia-smi --query-gpu=index,name,power.draw,power.limit,temperature.gpu --format=csv -l 5

# cap each GPU so the whole chassis fits the contracted power budget
sudo nvidia-smi -pm 1          # persistence mode
sudo nvidia-smi -i 0 -pl 300   # example: limit GPU 0 to 300 W

# whole-server reading from the BMC (IPMI)
ipmitool -I lanplus -H 198.51.100.20 -U admin -P '***' dcmi power reading

Network: 10Gbps, your own IPs and BGP

Inference traffic is small per request but latency-sensitive, so prioritise clean routing and your own address space over raw bandwidth.

  • 10Gbps uplinks for model downloads, replication between sites and high-concurrency APIs.
  • Your own IP block: lease IPv4 from IMIDC (up to full /24) or bring your own.
  • BGP with your ASN: announce your prefix from Tokyo and Los Angeles, for example, so you can move traffic between sites without renumbering; IMIDC also offers Anycast.
  • Hong Kong carries CN2 GIA (AS4809) routing to mainland China, and Los Angeles offers Unicom 9929/4837 and CN2 routes — useful if users of a lawful self-hosted model are in mainland China (hosting outside the mainland needs no ICP filing, but content must still comply with applicable law).
  • Expose inference only through a gateway with authentication and rate limiting; keep the model server bound to localhost or a private VLAN.
# example: serve an open-weight model with vLLM, API on localhost only
docker run -d --name vllm --gpus all --restart unless-stopped \
  -p 127.0.0.1:8000:8000 --ipc=host \
  -v /data/models:/models \
  vllm/vllm-openai:latest \
  --model /models/your-model --tensor-parallel-size 4

curl -s http://127.0.0.1:8000/v1/models

Choosing a location by latency to users

Pick the city nearest to the majority of requests, then test from where your users actually are.

IMIDC locationTypical audienceNotes
Tokyo, JapanJapan, Korea, wider North AsiaNative Japanese IPs; strong regional peering
Hong KongGreater China, Southeast AsiaCN2 GIA to mainland China; check export-control rules for advanced GPUs
SingaporeSoutheast Asia, India, OceaniaConfigured via sales
Los Angeles, USANorth America, trans-PacificUnicom 9929/4837 and CN2 routes to China
# from a user-side test host, compare candidate locations
mtr -rwc 100 203.0.113.10     # Tokyo test IP (example)
mtr -rwc 100 203.0.113.20     # Singapore test IP (example)

Export controls: read this before shipping GPUs

Advanced AI chips are subject to export-control rules, and you, as owner and shipper, are responsible for compliance.

  • US export rules (EAR) restrict certain advanced computing chips and systems for some destinations and end users; several of those restrictions also apply to Hong Kong and Macau.
  • Japan, Singapore and other jurisdictions have their own export and re-export rules.
  • Check the classification of your exact GPU model with your vendor, document the end use and end user, and obtain any required licenses before shipping.
  • IMIDC provides space, power, network and remote hands; it does not assess or take responsibility for the export status of customer hardware.

This is a brief overview, not legal advice. Consult a trade-compliance specialist.

Shipping, customs and install checklist

Most delays come from paperwork and missing parts, not from racking.

  1. Open a ticket with IMIDC first: confirm location, rack units, power (kW and feeds), delivery address and receiving hours.
  2. Prepare a commercial invoice with model, serial numbers, HS codes and declared value; agree who acts as importer of record.
  3. Insure the shipment; use original or foam-fitted packaging and send GPUs installed or in anti-static packaging.
  4. Label every box with your company name and IMIDC ticket number.
  5. Include rails, power cords (C13/C19 as needed), optics or DAC cables, and spare disks/fans.
  6. Pre-configure the BMC/IPMI with a static IP and a strong password, and set BIOS power-restore to "on".
  7. Send a rack diagram and cabling plan; remote hands install, power on and confirm link and BMC access.

After install, remote hands can handle GPU or disk swaps from your spares, reseats and power cycles via ticket.

Which IMIDC setup fits

Match the colocation size to measured power, then add network services.

  • Single inference box (1-4 GPUs): 1U-4U colocation in Tokyo or Los Angeles, plus a few IPs.
  • 8-GPU server: a full rack with dedicated power, discussed with sales before purchase.
  • Multi-region API: racks in Tokyo and Los Angeles announcing your own prefix via BGP, with Anycast if needed.
  • CPU-side services: put the API gateway, vector DB or queue on an IMIDC dedicated server in the same data center.

High-density power, Singapore colocation and multi-rack projects are custom configurations quoted by IMIDC sales.

FAQ

Where can I colocate my own GPU servers in Tokyo or Hong Kong?

IMIDC offers colocation from 1U to a full rack in Tokyo, Hong Kong, Singapore and Los Angeles, with remote hands and 10Gbps uplinks. Confirm your power requirement with sales before shipping hardware.

Does IMIDC rent GPU servers?

No. IMIDC provides colocation space, power, network and IP resources for hardware you own. You can pair colocated GPUs with IMIDC dedicated servers for CPU workloads.

How much power does an 8-GPU server need in colocation?

Typically several kW and sometimes over 10 kW at full load, depending on the GPU model. Measure under real load, consider power caps, and plan a full rack with sales.

Can I announce my own IP block from colocated GPU servers?

Yes. IMIDC supports BGP announcement with your own ASN, IPv4 leasing including full /24 blocks, and Anycast.

Planning a GPU deployment? Review IMIDC colocation and data centers, BGP and Anycast and IP resources, then contact sales with your server model and power figures, or open a ticket to schedule an install.

Was this answer helpful?

Related Tutorials