MS Mukul Sharma
Resume
vLLM & RAG Infrastructure // CASE STUDY ARCHITECTURE

Vector Search & GPU Inference Infrastructure

High-performance vLLM GPU cluster, vector embedding pipelines, AWS-to-GCP modernization, and hybrid FastAPI + Node.js services.

vLLM GPU Serving Vector Database RAG Pipeline AWS to GCP Migration FastAPI / Node.js

01 // Problem Statement

Third-party LLM API costs and latency bottlenecks limited custom model fine-tuning and domain-specific RAG accuracy for enterprise customer search requirements.

02 // Technical Constraints

  • Sub-100ms vector retrieval target [ADD: actual vector DB latency p95 ms].
  • Zero downtime during AWS-to-GCP cloud infrastructure migration.
  • Efficient GPU memory utilization using vLLM PagedAttention.

Executive Summary

To power real-time AI capabilities across our CRM platform, relying solely on public LLM endpoints introduced latency spikes and cost unpredictability.

I architected our core AI vector search and inference infrastructure, modernizing our stack from AWS to GCP and deploying self-hosted vLLM GPU clusters paired with scalable vector databases.

System Architecture

[Node.js Business Services] ──(gRPC / Internal REST)──> [FastAPI AI Services]
                                                               │
                                       ┌───────────────────────┴───────────────────────┐
                                       ▼                                               ▼
                         [vLLM GPU Inference Cluster]                [Vector Database Cluster]
                         (GCP Compute Engine / NVIDIA)               (Embeddings & Hybrid Search)

Key Technical Achievements

  1. AWS-to-GCP Modernization: Led the complete cloud infrastructure migration from AWS to GCP, re-architecting service deployment templates, secrets management, and network security.
  2. vLLM Optimization: Leveraged vLLM’s PagedAttention mechanism for high-throughput batch inference, serving custom open LLMs for structured parsing tasks.
  3. Hybrid Service Layer: Structured a high-performance polyglot architecture combining Node.js microservices for core web APIs with FastAPI Python services for AI tensor workflows.

03 // Key Decisions & Trade-offs

vLLM GPU Serving vs Cloud Managed Inference

Deployed dedicated vLLM GPU inference instances on GCP, cutting per-token inference overhead [ADD: % cost reduction] and giving full control over context window lengths.

Hybrid Microservices Architecture (FastAPI + Node.js)

Used Node.js for high-concurrency CRM business logic and FastAPI for heavy AI/vector computations, communicating via gRPC / HTTP internal channels.

04 // Measurable Results

  • Successful migration of core backend infrastructure from AWS to GCP with zero customer outage.
  • Production RAG vector search supporting real-time document grounding.
  • High throughput LLM inference powered by optimized GPU serving.
  • AWS-to-GCP cloud modernization improving infrastructure cost efficiency.

05 // What I'd Do Next

  • Implement dynamic GPU auto-scaling based on token queue depth.
  • Add speculative decoding to accelerate token generation latency.