Start typing to search across invoices, services, domains, tickets, and more...
The most practical way to run a private AI knowledge base (RAG) for customer support is to keep your documents, vector database and chat logs on a dedicated server you control, and call an official LLM API only for embeddings and answers. IMIDC provides dedicated servers in Tokyo, Japan; Singapore; Taipei, Taiwan and Los Angeles, USA — all regions where the OpenAI, Anthropic Claude and Google Gemini APIs are officially available — so a cross-border business can keep customer data in the region it serves.
Retrieval-augmented generation (RAG) answers customer questions from your own documents instead of the model's memory, which cuts hallucinations and keeps answers on-policy.
The flow is simple: your manuals, FAQs, return policies, product sheets and past ticket resolutions are split into chunks, each chunk is turned into a vector (an embedding), and the vectors are stored in a vector database. When a customer or agent asks a question, the system embeds the question, retrieves the closest chunks, and sends only those chunks plus the question to the LLM, which writes an answer with citations.
For a cross-border seller handling Japanese, English and Chinese tickets, one multilingual index can serve every channel.
A dedicated server wins when you have many documents, strict data rules or steady query volume, because RAM, disk and cost are fixed and the data never leaves your box except for the chunks you choose to send.
| Factor | SaaS RAG chatbot | Self-hosted RAG on an IMIDC dedicated server |
|---|---|---|
| Where documents and chat logs live | Vendor's cloud, region often unclear | Your server in Tokyo, Singapore, Taipei or Los Angeles |
| Pricing model | Per seat, per message or per document | Fixed monthly server fee + pay-per-token API usage |
| Vector capacity | Plan limits | Bounded only by your RAM and NVMe |
| Access control | Vendor's model | Your own ACLs per team, brand or customer tier |
| Model choice | Usually one vendor | Switch between OpenAI, Claude and Gemini, or a local embedding model |
| Effort | Low | You run Docker, backups and updates |
Memory is the key resource. One million chunks with 1536-dimension float32 embeddings need roughly 6 GB of raw vector data, plus index overhead; Qdrant and pgvector both perform best when the index sits in RAM. A dedicated server with 64-128 GB of RAM leaves room for vectors, PostgreSQL, the ingestion worker and caching without the noisy-neighbour effects of a small VPS.
Put the knowledge base where your customers and your data obligations are, and only in a region where the AI APIs are officially supported.
| IMIDC location | Best for | AI API availability | Data-protection framework (brief) |
|---|---|---|---|
| Tokyo, Japan | Japanese customers, Rakuten/Amazon JP sellers | Official | APPI |
| Singapore | Southeast Asia, multi-country support desks | Official | PDPA |
| Taipei, Taiwan | Traditional Chinese support, Taiwan market | Official | Taiwan PDPA |
| Los Angeles, USA | North American customers, US marketplaces | Official | US state privacy laws |
| Hong Kong / Moscow, Russia | Not recommended for this workload | Not officially available | — |
Three containers are enough to start: a vector database, PostgreSQL for metadata and chat history, and your own ingestion/chat API.
Install Docker first (see our Docker installation guide in the help center), then create the following files. Every port is bound to 127.0.0.1 and published to the internet only through Nginx with HTTPS.
# docker-compose.yml - private RAG stack (all ports bound to localhost)
services:
qdrant:
image: qdrant/qdrant:v1.12.4
restart: unless-stopped
volumes:
- ./qdrant_storage:/qdrant/storage
ports:
- "127.0.0.1:6333:6333"
environment:
QDRANT__SERVICE__API_KEY: ${QDRANT_API_KEY}
postgres:
image: pgvector/pgvector:pg16
restart: unless-stopped
environment:
POSTGRES_USER: rag
POSTGRES_PASSWORD: ${PG_PASSWORD}
POSTGRES_DB: kb
volumes:
- ./pgdata:/var/lib/postgresql/data
ports:
- "127.0.0.1:5432:5432"
rag-api:
build: ./rag-api # your ingestion + retrieval + chat service
restart: unless-stopped
env_file: .env # OPENAI_API_KEY / ANTHROPIC_API_KEY / GEMINI_API_KEY
depends_on: [qdrant, postgres]
ports:
- "127.0.0.1:8080:8080" # publish through Nginx + HTTPS only
# .env (chmod 600, never commit)
QDRANT_API_KEY=change-me-long-random
PG_PASSWORD=change-me-long-random
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
EMBED_MODEL=text-embedding-3-small
CHUNK_TOKENS=500
CHUNK_OVERLAP=60
If you prefer one database, skip Qdrant and use pgvector for both metadata and vectors:
-- pgvector alternative: one table, one HNSW index
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (
id bigserial PRIMARY KEY,
doc_id text NOT NULL,
lang text,
acl text[], -- which teams/customers may see it
content text NOT NULL,
embedding vector(1536)
);
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);
Qdrant is the better choice above a few million vectors or when you need payload filtering and quantization; pgvector is simpler when your team already runs PostgreSQL. Verify outbound API access from the server before ingestion:
# confirm the API endpoints are reachable from this server
curl -s -o /dev/null -w "%{http_code}\n" https://api.openai.com/v1/models \
-H "Authorization: Bearer $OPENAI_API_KEY"
curl -s -o /dev/null -w "%{http_code}\n" https://api.anthropic.com/v1/models \
-H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01"
Answer quality depends more on chunking and metadata than on the model.
Hosting the database in-region helps, but any text you send to an LLM API is processed by that provider, so minimise and document what leaves the server.
This is general information, not legal advice. Confirm your obligations with qualified counsel.
Size the server to your vector count and concurrency, then pick the region closest to your customers.
Singapore servers and custom RAM/disk configurations are arranged through IMIDC sales. Moving from another provider? IMIDC offers free server migration.
IMIDC offers dedicated servers in Tokyo and Singapore with native local IPs, where the OpenAI, Claude and Gemini APIs are officially available. You keep the vector database and documents on your own hardware and pay a fixed monthly fee.
Not with the OpenAI, Claude or Gemini APIs: they are not officially available in Hong Kong, mainland China or Russia. Choose Tokyo, Singapore, Taipei or Los Angeles for API-based AI workloads.
Roughly 6 GB per million 1536-dimension vectors before index overhead, so plan 2x that for HNSW and the rest of the stack. Scalar quantization in Qdrant can cut vector memory by about four times.
No hosting choice alone makes you compliant. In-region storage simplifies things, but you still need lawful purposes, notices and safeguards for any data sent to API providers; this is not legal advice.
Ready to build your support knowledge base? Compare Tokyo dedicated servers, Taipei dedicated servers and Los Angeles dedicated servers, or contact IMIDC sales for Singapore and custom RAM configurations. Existing clients can open a ticket for sizing advice.