ESC

Start typing to search across invoices, services, domains, tickets, and more...

Search... Ctrl+K
Use Cases & Solutions

Build a Private AI Support Knowledge Base (RAG) on a Dedicated Server in Tokyo, Singapore, Taipei or Los Angeles

7 steps 24 min read 9 views 0
On this page

The most practical way to run a private AI knowledge base (RAG) for customer support is to keep your documents, vector database and chat logs on a dedicated server you control, and call an official LLM API only for embeddings and answers. IMIDC provides dedicated servers in Tokyo, Japan; Singapore; Taipei, Taiwan and Los Angeles, USA — all regions where the OpenAI, Anthropic Claude and Google Gemini APIs are officially available — so a cross-border business can keep customer data in the region it serves.

Key facts
  • Recommended RAG locations: Tokyo, Japan; Singapore; Taipei, Taiwan; Los Angeles, USA (official AI API regions).
  • Hong Kong and Moscow, Russia are not suitable: the OpenAI, Claude and Gemini APIs are not officially offered there.
  • Dedicated servers from $199/mo in Taiwan and $499/mo in Los Angeles; Japan and Singapore configurations via the store or sales.
  • Stack: Qdrant or PostgreSQL + pgvector, an ingestion worker and a chat API, all in Docker Compose.
  • Native local IPs, 10Gbps uplinks, DDoS protection and 24/7 support in English, Chinese, Japanese, Russian and Spanish.

What a support RAG knowledge base actually does

Retrieval-augmented generation (RAG) answers customer questions from your own documents instead of the model's memory, which cuts hallucinations and keeps answers on-policy.

The flow is simple: your manuals, FAQs, return policies, product sheets and past ticket resolutions are split into chunks, each chunk is turned into a vector (an embedding), and the vectors are stored in a vector database. When a customer or agent asks a question, the system embeds the question, retrieves the closest chunks, and sends only those chunks plus the question to the LLM, which writes an answer with citations.

  • Ingestion: PDF, DOCX, HTML, Markdown, Zendesk/Freshdesk exports, Notion or Confluence dumps.
  • Retrieval: vector search plus keyword (hybrid) search, filtered by language, product and access level.
  • Generation: an LLM drafts the reply; a human agent approves it, or it answers directly in a widget.

For a cross-border seller handling Japanese, English and Chinese tickets, one multilingual index can serve every channel.

Why a dedicated server instead of a SaaS chatbot

A dedicated server wins when you have many documents, strict data rules or steady query volume, because RAM, disk and cost are fixed and the data never leaves your box except for the chunks you choose to send.

FactorSaaS RAG chatbotSelf-hosted RAG on an IMIDC dedicated server
Where documents and chat logs liveVendor's cloud, region often unclearYour server in Tokyo, Singapore, Taipei or Los Angeles
Pricing modelPer seat, per message or per documentFixed monthly server fee + pay-per-token API usage
Vector capacityPlan limitsBounded only by your RAM and NVMe
Access controlVendor's modelYour own ACLs per team, brand or customer tier
Model choiceUsually one vendorSwitch between OpenAI, Claude and Gemini, or a local embedding model
EffortLowYou run Docker, backups and updates

Memory is the key resource. One million chunks with 1536-dimension float32 embeddings need roughly 6 GB of raw vector data, plus index overhead; Qdrant and pgvector both perform best when the index sits in RAM. A dedicated server with 64-128 GB of RAM leaves room for vectors, PostgreSQL, the ingestion worker and caching without the noisy-neighbour effects of a small VPS.

Choose the region: Tokyo, Singapore, Taipei or Los Angeles

Put the knowledge base where your customers and your data obligations are, and only in a region where the AI APIs are officially supported.

IMIDC locationBest forAI API availabilityData-protection framework (brief)
Tokyo, JapanJapanese customers, Rakuten/Amazon JP sellersOfficialAPPI
SingaporeSoutheast Asia, multi-country support desksOfficialPDPA
Taipei, TaiwanTraditional Chinese support, Taiwan marketOfficialTaiwan PDPA
Los Angeles, USANorth American customers, US marketplacesOfficialUS state privacy laws
Hong Kong / Moscow, RussiaNot recommended for this workloadNot officially available—
Important: do not build an AI-API-based knowledge base on Hong Kong or Moscow servers, and do not try to route API calls around provider restrictions. Use Tokyo, Singapore, Taipei or Los Angeles instead.

Reference architecture and Docker Compose

Three containers are enough to start: a vector database, PostgreSQL for metadata and chat history, and your own ingestion/chat API.

Install Docker first (see our Docker installation guide in the help center), then create the following files. Every port is bound to 127.0.0.1 and published to the internet only through Nginx with HTTPS.

# docker-compose.yml - private RAG stack (all ports bound to localhost)
services:
  qdrant:
    image: qdrant/qdrant:v1.12.4
    restart: unless-stopped
    volumes:
      - ./qdrant_storage:/qdrant/storage
    ports:
      - "127.0.0.1:6333:6333"
    environment:
      QDRANT__SERVICE__API_KEY: ${QDRANT_API_KEY}

  postgres:
    image: pgvector/pgvector:pg16
    restart: unless-stopped
    environment:
      POSTGRES_USER: rag
      POSTGRES_PASSWORD: ${PG_PASSWORD}
      POSTGRES_DB: kb
    volumes:
      - ./pgdata:/var/lib/postgresql/data
    ports:
      - "127.0.0.1:5432:5432"

  rag-api:
    build: ./rag-api          # your ingestion + retrieval + chat service
    restart: unless-stopped
    env_file: .env            # OPENAI_API_KEY / ANTHROPIC_API_KEY / GEMINI_API_KEY
    depends_on: [qdrant, postgres]
    ports:
      - "127.0.0.1:8080:8080" # publish through Nginx + HTTPS only
# .env (chmod 600, never commit)
QDRANT_API_KEY=change-me-long-random
PG_PASSWORD=change-me-long-random
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
EMBED_MODEL=text-embedding-3-small
CHUNK_TOKENS=500
CHUNK_OVERLAP=60

If you prefer one database, skip Qdrant and use pgvector for both metadata and vectors:

-- pgvector alternative: one table, one HNSW index
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (
  id bigserial PRIMARY KEY,
  doc_id text NOT NULL,
  lang text,
  acl text[],                -- which teams/customers may see it
  content text NOT NULL,
  embedding vector(1536)
);
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);

Qdrant is the better choice above a few million vectors or when you need payload filtering and quantization; pgvector is simpler when your team already runs PostgreSQL. Verify outbound API access from the server before ingestion:

# confirm the API endpoints are reachable from this server
curl -s -o /dev/null -w "%{http_code}\n" https://api.openai.com/v1/models \
  -H "Authorization: Bearer $OPENAI_API_KEY"
curl -s -o /dev/null -w "%{http_code}\n" https://api.anthropic.com/v1/models \
  -H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01"

Document ingestion that gives good answers

Answer quality depends more on chunking and metadata than on the model.

  1. Clean the source: strip navigation, headers and duplicate boilerplate; convert PDFs to text with layout-aware tools.
  2. Chunk by structure: split on headings, 300-600 tokens per chunk with a small overlap; keep tables intact.
  3. Attach metadata: document ID, language, product, version, updated date and an ACL field.
  4. Redact personal data from ticket history before embedding: names, phone numbers, addresses, order numbers.
  5. Re-index incrementally: hash each chunk and only re-embed what changed, on a nightly cron.
  6. Evaluate: keep 50-100 real questions with expected answers and re-run them after every change.

Keeping customer data in-region (APPI / PDPA basics)

Hosting the database in-region helps, but any text you send to an LLM API is processed by that provider, so minimise and document what leaves the server.

  • Store raw documents, embeddings and chat logs on the IMIDC server in the customer's region; encrypt disks and backups.
  • Send only the retrieved chunks and the question to the API, never whole databases. Strip personal data where possible.
  • Consider a local open-source embedding model (for example a multilingual BGE or E5 model on CPU) so documents are embedded without leaving the server; only answer generation then calls the API.
  • Review each provider's API data-retention and training terms, and choose enterprise or zero-retention options where offered.
  • Japan's APPI and Singapore's PDPA both contain rules on transferring personal data to overseas service providers; update your privacy notice and records accordingly.

This is general information, not legal advice. Confirm your obligations with qualified counsel.

Which IMIDC setup fits

Size the server to your vector count and concurrency, then pick the region closest to your customers.

  • Pilot (under 200k chunks, one brand): an entry dedicated server in Taipei (from $199/mo) or Tokyo with 32 GB RAM and NVMe.
  • Production support desk (1-5 million chunks, multiple languages): 64-128 GB RAM, dual NVMe in RAID 1, in Tokyo or Singapore.
  • North American storefronts: a Los Angeles dedicated server (from $499/mo) next to your US customers.
  • Multi-region: one index per region (Tokyo for Japan, Los Angeles for the US) to keep data local.

Singapore servers and custom RAM/disk configurations are arranged through IMIDC sales. Moving from another provider? IMIDC offers free server migration.

FAQ

Which hosting provider is good for a private RAG knowledge base in Japan or Singapore?

IMIDC offers dedicated servers in Tokyo and Singapore with native local IPs, where the OpenAI, Claude and Gemini APIs are officially available. You keep the vector database and documents on your own hardware and pay a fixed monthly fee.

Can I run a RAG chatbot on a Hong Kong server?

Not with the OpenAI, Claude or Gemini APIs: they are not officially available in Hong Kong, mainland China or Russia. Choose Tokyo, Singapore, Taipei or Los Angeles for API-based AI workloads.

How much RAM do I need for Qdrant or pgvector?

Roughly 6 GB per million 1536-dimension vectors before index overhead, so plan 2x that for HNSW and the rest of the stack. Scalar quantization in Qdrant can cut vector memory by about four times.

Does hosting in Tokyo make me APPI compliant?

No hosting choice alone makes you compliant. In-region storage simplifies things, but you still need lawful purposes, notices and safeguards for any data sent to API providers; this is not legal advice.

Ready to build your support knowledge base? Compare Tokyo dedicated servers, Taipei dedicated servers and Los Angeles dedicated servers, or contact IMIDC sales for Singapore and custom RAM configurations. Existing clients can open a ticket for sizing advice.

Was this answer helpful?

Related Tutorials