Vector Search & GPU Inference Infrastructure
High-performance vLLM GPU cluster, vector embedding pipelines, AWS-to-GCP modernization, and hybrid FastAPI + Node.js services.
01 // Problem Statement
Third-party LLM API costs and latency bottlenecks limited custom model fine-tuning and domain-specific RAG accuracy for enterprise customer search requirements.
02 // Technical Constraints
- Sub-100ms vector retrieval target [ADD: actual vector DB latency p95 ms].
- Zero downtime during AWS-to-GCP cloud infrastructure migration.
- Efficient GPU memory utilization using vLLM PagedAttention.
Executive Summary
To power real-time AI capabilities across our CRM platform, relying solely on public LLM endpoints introduced latency spikes and cost unpredictability.
I architected our core AI vector search and inference infrastructure, modernizing our stack from AWS to GCP and deploying self-hosted vLLM GPU clusters paired with scalable vector databases.
System Architecture
[Node.js Business Services] ──(gRPC / Internal REST)──> [FastAPI AI Services]
│
┌───────────────────────┴───────────────────────┐
▼ ▼
[vLLM GPU Inference Cluster] [Vector Database Cluster]
(GCP Compute Engine / NVIDIA) (Embeddings & Hybrid Search)
Key Technical Achievements
- AWS-to-GCP Modernization: Led the complete cloud infrastructure migration from AWS to GCP, re-architecting service deployment templates, secrets management, and network security.
- vLLM Optimization: Leveraged vLLM’s PagedAttention mechanism for high-throughput batch inference, serving custom open LLMs for structured parsing tasks.
- Hybrid Service Layer: Structured a high-performance polyglot architecture combining Node.js microservices for core web APIs with FastAPI Python services for AI tensor workflows.
03 // Key Decisions & Trade-offs
vLLM GPU Serving vs Cloud Managed Inference
Deployed dedicated vLLM GPU inference instances on GCP, cutting per-token inference overhead [ADD: % cost reduction] and giving full control over context window lengths.
Hybrid Microservices Architecture (FastAPI + Node.js)
Used Node.js for high-concurrency CRM business logic and FastAPI for heavy AI/vector computations, communicating via gRPC / HTTP internal channels.
04 // Measurable Results
- Successful migration of core backend infrastructure from AWS to GCP with zero customer outage.
- Production RAG vector search supporting real-time document grounding.
- High throughput LLM inference powered by optimized GPU serving.
- AWS-to-GCP cloud modernization improving infrastructure cost efficiency.
05 // What I'd Do Next
- Implement dynamic GPU auto-scaling based on token queue depth.
- Add speculative decoding to accelerate token generation latency.