Flagship production systems
End-to-end, high-availability architectures engineered and battle-tested in live enterprise environments across retail, telecom, forestry, insurance, enterprise AI, and banking.
Production-grade AI systems, recommendation platforms, MLOps, and enterprise software engineering.
I have designed production-grade AI platforms and high-throughput ML infrastructure built for enterprise-scale operations. By combining low-latency recommendation engines, agentic LLMs, and robust MLOps pipelines, I eliminate system bottlenecks and ensure sub-second inference at scale. My primary focus is turning complex data ecosystems into fault-tolerant decision engines that drive multi-million-dollar business ROI.
By the numbers
A concise view of the engineering surface represented by this showcase. The numbers describe the systems on this page, not INVENTED CAREER statistics.
End-to-end, high-availability architectures engineered and battle-tested in live enterprise environments across retail, telecom, forestry, insurance, enterprise AI, and banking.
Bespoke algorithmic solutions designed to convert complex domain challenges into direct bottom-line ROI, sub-second latency reduction, and high-throughput decision automation.
Deep architectural mastery spanning multi-stage recommendation engines, real-time Computer Vision, Enterprise LLMs, automated Document AI, and high-precision predictive modeling.
Creator and maintainer of cooprecsys: a highly modular, production-ready Python framework designed to streamline candidate retrieval and hybrid recommendation engineering.
Why companies hire me
Six engineering behaviors shape every system in this showcase.
Start with the operational decision, measurable outcome, and failure cost, not the model.
Treat data, inference, deployment, observability, testing, and ownership as one system.
Design boundaries, interfaces, lineage, and trade-offs so systems can evolve without becoming fragile.
Make experimentation reproducible and production behavior observable from training through inference.
Build dependable feature and data workflows that keep model development aligned with operational reality.
Translate architecture and uncertainty into decisions that engineering, product, and leadership can act on.
Visual Executive Summary
A production-grade recommendation platform built on reliable feature engineering, scalable freshness, ML reproducibility, operational reliability, and measurable outcomes.
p99 end-to-end response time achieved via ONNX Runtime & C++ optimization
NDCG@10 score achieved by LightGBM LambdaRank on holdout evaluation sets
Sub-millisecond Redis online feature retrieval during candidate scoring
Recommended
Top Match
Trending
New Arrival Decoupled architecture between candidate generation and re-ranking enables sub-millisecond processing across millions of catalog items.
Real-time data integrity powered by a Feature Store ensures instantaneous reflection of stock availability and user interaction intent.
Engineered to systematically bridge offline ML experimentation with low-latency, highly reliable online serving.
Centralized Grafana observability to concurrently monitor technical infrastructure performance and core business metrics.
Automated testing and promotion workflows for ML models transitioning from staging to production.
High-availability protections against downstream service failures.
A production-grade recommendation system goes far beyond building ML models. Sustained competitive advantage stems from strict feature contracts, repeatable workflows, automated model promotion, end-to-end observability, and feedback loops directly wired to real business KPIs.
Interactive KMS Topology Explorer
Trace the end-to-end data path across multimodal document ingestion, async enrichment, hybrid vector-graph retrieval, and real-time observability.
Engineered for strict operational boundaries and sub-100ms retrieval SLAs, this interactive topology maps every production subsystem. Select any node to inspect its execution context, state management, and underlying trade-offs.
Hydrating topology canvas...
Hydrating topology canvas...
Production AI systems
The flagship recommendation platform is only one part of the engineering story. These systems show how the same production discipline adapts across forestry, insurance, telecommunications, document intelligence, and banking.
Production AI system 02
An enterprise-grade computer vision and geospatial intelligence platform running in AWS VPC, engineered to process drone and satellite payloads into real-time operational insights for field workers and executive management.
Technology stack
Constraints
System Benchmarks & SLAs
Benchmarked on AWS VPC GPU clusters, spatial tiling with Redis caching delivers sub-110ms tile inference and 99.95% offline sync reliability for high-concurrency drone and field crew operations.
Deployment
Monitoring
Business impact
Lessons learned
ADR-001
Alternatives
Single Generalist Object Detector · Traditional Remote Sensing Indexing (NDVI)
Decision
Deploy Vision Transformers for macro classification (weeds, flooding, wild trees) and combine DINOv3 with Detectron2 for micro object detection (young tree counting and early disease identification).
Business impact
Maximizes detection sensitivity across varying growth stages while maintaining computational efficiency inside AWS VPC GPU clusters.
ADR-002
Alternatives
Direct Database Spatial Queries · Static GeoJSON File Exports
Decision
Persist relational spatial lineage in MS SQL Server while serving dynamic prediction states and boundary shapefiles via an in-memory Redis caching tier.
Business impact
Slashes API response latency for hundreds of simultaneous field workers and drone pilots uploading operational data in real time.
Production AI/ML Architecture 03
An end-to-end, highly scalable Agentic ML & Distributed Systems platform built for top-tier OTAs (e.g., Traveloka / Tiket.com). Features hybrid multi-tier GraphRAG (ColBERTv2 + Neo4j Cypher), Small Language Model (SLM) intent routing, vLLM constrained decoding via Context-Free Grammars (CFG), Temporal.io Saga orchestration for multi-GDS ACID transactions, and an Active Learning DPO feedback loop.
Technology stack
Constraints
System Benchmarks & SLAs
Evaluated under a simulated 100k RPM traffic spike, the system achieved an 82.4% autonomous FCR with an average TTFT of 180ms. Semantic caching intercepted 28.5% of incoming queries with < 12ms response time.
Deployment
Monitoring
Business impact
Lessons learned
ADR-001
Alternatives
Single-Prompt LLM (Claude 3.5 / Llama 3.3 70B) for All Tasks · Static Regex/Rule Matcher
Decision
Deploy a fine-tuned Qwen 2.5-0.5B Small Language Model (SLM) for intent classification & routing, paired with RedisVL Vector Semantic Caching ahead of the main LLM orchestrator.
Business impact
Intersects ~28% of repetitive queries under 12ms latency and cuts overall inference infrastructure operational costs by 65%.
ADR-002
Alternatives
Standard Dense Embedding Vector Search (OpenAI / BGE-M3) · Hardcoded Business Rules Engine
Decision
Combine Neo4j Cypher Graph Queries for deterministic structural relations (Booking ID -> PNR -> Airline Rules) with Qdrant ColBERTv2 multi-vector search, re-ranked via Cross-Encoder (BGE-Reranker-Large).
Business impact
Guarantees 100% accuracy on complex multi-hop refund and reschedule calculations across 50+ partner airlines under PM 89/2015.
ADR-003
Alternatives
Pure Python ReAct Loops with Native Asyncio · Celery Distributed Task Queue
Decision
Use LangGraph for LLM multi-agent reasoning and plan generation, but delegate state execution to Temporal.io Saga Workflows enforcing Try-Confirm-Cancel (TCC) transactional patterns.
Business impact
Prevents partial state failures (e.g., ticket cancelled on airline side but refund failed on OTA wallet side) with automated multi-system rollbacks.
ADR-004
Alternatives
Standard OpenAI-Style Native Function Calling · Unconstrained Prompt-Based JSON Output Parsing
Decision
Enforce strict Context-Free Grammar (CFG) constrained decoding at the GPU token sampling level via SGLang / Outlines on vLLM, paired with NeMo Guardrails and Llama Guard 3 security middleware.
Business impact
Eliminates 100% of JSON syntax/schema parsing errors and halts prompt injection attempts before non-deterministic LLM output can hit critical financial APIs (OTA Wallet & GDS write endpoints).
Production AI System 04
An enterprise-grade Customer Data Platform (CDP) and AI intelligence layer unifying distributed multi-dealer databases across Southeast Asia into a single identity-resolved customer graph for hyper-personalized buyer targeting, social identity enrichment, and high-intent lead routing.
Technology stack
Constraints
System Benchmarks & SLAs
Streaming across 200+ regional dealerships, probabilistic identity resolution and Feast Feature Store deliver 98.4% match accuracy and sub-18ms ML feature retrieval to drive sales lead conversions.
Deployment
Monitoring
Business impact
Lessons learned
ADR-001
Alternatives
Single Unique Email Matching · Centralized Master DB Replication
Decision
Implement a hybrid identity graph (Neo4j) pairing exact matches (phone/hashed ID) with probabilistic fuzzy matching (name, address, social handles).
Business impact
Successfully linked 72% of previously disconnected offline showroom visitors to online behavioral profiles without exposing PII.
ADR-002
Alternatives
Batch Nightly ETL to Local CRM · Direct API Queries to Legacy DMS
Decision
Deploy Feast Feature Store backed by Redis to serve offline-trained propensity features directly to real-time scoring endpoints.
Business impact
Reduced feature retrieval latency from 4.2 seconds to 12 milliseconds, enabling instant sales rep alerts during active buyer browsing.
Production AI system 05
A document intelligence pipeline that converts heterogeneous business documents into validated structured data with confidence-aware human review.
Technology stack
Constraints
System Benchmarks & SLAs
Combining hybrid OCR and multimodal AI, the document pipeline achieves 98.2% extraction precision and an 86.0% straight-through processing rate with automated failover to human review queues.
Deployment
Monitoring
Business impact
Lessons learned
ADR-001
Alternatives
Single multimodal prompt · OCR-only extraction
Decision
Use explicit OCR, layout, classification, and reasoning stages so each failure mode can be observed and controlled.
Business impact
Improves debuggability and allows quality gates to be placed before structured data reaches business systems.
ADR-002
Alternatives
Fully automated extraction · Manual processing for every document
Decision
Route low-confidence or validation-failing outputs to a human review path instead of silently publishing them.
Business impact
Balances automation efficiency with the trust requirements of enterprise data workflows.
Production AI system 06
A supervised prediction platform that converts core-banking data into calibrated payoff propensity signals for CRM and business decision workflows.
Technology stack
Constraints
System Benchmarks & SLAs
Processing 1.2M banking records per minute, the calibrated LightGBM model achieves a 0.884 ROC-AUC with TreeSHAP explainability, serving real-time propensity scores to CRM workflows in under 35ms.
Deployment
Monitoring
Business impact
Lessons learned
ADR-001
Alternatives
Logistic regression · XGBoost · Deep neural network
Decision
Use LightGBM when structured tabular features and fast retraining provide a strong fit for the prediction problem.
Business impact
Keeps iteration and serving complexity proportional to the structured-data problem while retaining useful feature attribution workflows.
ADR-002
Alternatives
Raw model scores only · Threshold-only monitoring
Decision
Monitor calibration and score distributions alongside conventional model performance signals.
Business impact
Makes downstream decision thresholds more trustworthy as customer populations and data distributions change.
Open Source Engineering
An enterprise-grade hybrid recommendation framework engineered with high-performance Cython/OpenMP C-extensions, DuckDB columnar feature pipelines, Group-aware Learning-to-Rank (LTR), MLflow tracking, and SHAP explainability.
High-throughput collaborative filtering with parallel C-extensions via Cython & OpenMP for ultra-low latency matrix factorization.
Group-aware LambdaMART ranking implementation optimized for multi-objective CTR & conversion prediction.
In-memory analytical querying on SQL layers for zero-copy feature transform and model experiment tracking.
High-performance Two-Tower retrieval algorithm optimized with Cython prange multi-threading for ultra-low latency feature attribution and NDCG slice diagnostics.
Production Workflow
Engineering Roadmap
Engineering philosophy
Frameworks change. Models change. Infrastructure changes. The operating principles should remain stable.
A result that cannot be recreated is not a dependable production asset. Version data, features, code, configuration, and model artifacts.
Production AI needs technical telemetry and business signals. Latency, failures, drift, confidence, and outcomes belong in the same operating conversation.
Scale the architecture where the workload actually grows. Avoid premature infrastructure while keeping boundaries ready for the next stage.
Automate repeatable work across testing, training, deployment, validation, and monitoring so engineering attention stays on high-value decisions.
A technically elegant system is still a failure if it does not improve a real business decision. Start from the workflow and work backward to the technology.
Delivery process
A disciplined delivery path keeps discovery, engineering, and operational ownership connected.
Clarify the business decision, users, constraints, data reality, and definition of success.
Establish system boundaries, data flow, interfaces, deployment shape, and technical trade-offs.
Prove the highest-risk assumptions with the smallest credible end-to-end slice.
Harden the workflow with testing, versioning, deployment automation, security, and operational ownership.
Measure system health, model behavior, data quality, drift, and business outcomes continuously.
Document decisions, operating procedures, architecture, and ownership so the system outlives the implementation team.
Enterprise AI systems
If the challenge involves turning AI capability into a dependable production platform, let's discuss the architecture, constraints, and path to delivery.