RIHAD VARIAWA


Senior AI Engineer & Systems Architect

Architecting production-grade LLM pipelines, optimizing GPU inference (vLLM/TensorRT), and building resilient distributed systems for FinTech and Enterprise.

2,000+
AWS RPS

$2B+
Daily Volume

< 5ms
Exec Latency

About

MLOps & Systems Architect owning end-to-end model lifecycle infrastructure, zero-downtime brownfield migrations, and high-throughput data platforms processing $2B+ daily volume and 2 TB/hr streaming telemetry.Proven track record inheriting brittle, failing pipelines and stabilizing production through disciplined scope subtraction—replacing over-engineered multi-hop chains with deterministic rerankers, cutting p99 latency by 42% and cloud spend by 28%.I establish day-one production rigor: automated MLflow champion/challenger alias routing, Feast point-in-time feature stores, Great Expectations CI gates, and observable rollback paths across AWS and Azure (EKS/AKS).📍 Availability: Full Remote | USA (EST) Overlap & European (CET/GMT) AlignmentInterested? Get in touch!

E X P E R I E N C E

Independent Applied AI Lab
Principal AI Architect & Fractional CTO (B2B Contracting) | London, UK (2023 — Present)
Directed architecture and deployment of enterprise-scale AI/ML solutions and distributed pipelines.
- Information Retrieval & GraphRAG: Engineered advanced RAG architectures optimizing chunking, embeddings, and custom re-rankers for unstructured financial data.
- Applied ML & Anomaly Detection: Built GBDTs and anomaly detection pipelines handling 2 TB/hr telemetry on AWS.
- AI Observability & Guardrails: Designed comprehensive evaluation loops (LangSmith, Arize Phoenix) and secure AI gateways for regulated environments.


Systemiclogic
Senior Consultant Data Scientist / Data Engineer | South Africa (2020 — 2023)
Led deployment and optimization of scalable ML workflows and enterprise MLOps.
- Transactional Pipelines: Engineered real-time, audit-compliant async Python architectures blocking transactional anomalies for Tier-1 banking.
- Cloud Migration: Scaled stateful processing by 3x by migrating monolithic legacy databases to distributed AWS/Azure topologies.
- ML Integration: Primary technical partner safely integrating predictive risk models into core, strictly regulated banking workflows.


Venture-Backed FinTech / Digital Asset Broker
Senior Software Engineer (Platform Scaling & Transaction Routing) | London Area, UK (2015 - 2020)
Applied advanced analytics, ML, and stochastic processes to financial risk modeling.
- Elastic Architecture: Engineered an async AWS serverless pipeline (Lambda) seamlessly scaling to 2,000+ RPS and processing $2B+ daily volume.
- Immutability & Latency: Architected distributed DynamoDB topologies, crushing API timeout errors by 45%.


Historical Trading Desks & Banking Infrastructure
Foundational Engineering & Quantitative R&D | Shanghai & Seoul (2010 - 2015)
- Algorithmic Execution: Built ultra-low latency Python pricing engines handling $50M+ daily volume (sub-5 ms execution).
- Fault-Tolerant Infrastructure: Eliminated Single Points of Failure (SPOFs) ensuring 100% data ingestion uptime.

P R O J E C T S

Production DevSecOps & MLOps CI/CD Pipeline
Automated Quality Gates, Vault Secrets & MLflow Lifecycle Registry
Architected an enterprise 3-stage blocking release pipeline enforcing zero-trust credential extraction and automated data schema drift protection.
- Zero-Trust Secrets: Runtime HashiCorp Vault (KV v2) extraction, eliminating static credentials from Git commit history and build environments.
- Automated Quality Gate: Non-swallowing Great Expectations CI pipeline that automatically blocks malformed Pull Requests on schema drift.
- Model Lineage & Tagging: MLflow registry backed by PostgreSQL and SeaweedFS (S3), dynamically assigning @production aliases to live releases.
View Architecture & GitHub Repo ↗


High-Throughput Async ML Inference Engine
Decoupled Model Serving, Redis Task Broker & Polling Architecture
Architected a non-blocking model inference gateway that offloads compute-heavy predictions to background workers, dropping API response times to < 5ms.
- Async Gateway & Task Queue: Implemented HTTP 202 Accepted patterns with UUID task tokenization, eliminating web worker thread starvation under high concurrency.
- In-Memory Result Broker: Integrated Redis key-value storage (result:<task_id>) with automated TTL memory cleanup (ex=600) to prevent RAM leaks.
- Resilient Polling Protocol: Designed a dual-state GET /result/<task_id> endpoint for stateless result retrieval across distributed web nodes.
View Architecture & GitHub Repo ↗


Enterprise MLOps & Feature Store Architecture
Advanced MLOps & Feature Platform Validation
Executed end-to-end MLOps pipeline architectures incorporating Feature Stores, Model Registries, and Data Quality CI Gates.
- Feature Store: Feast Feature Stores (point-in-time joins, online SQLite materialization).
- Model Registry & Versioning: MLflow (champion/challenger alias routing, PyFunc wrappers) and DVC data versioning.
- Data Quality: Great Expectations automated data-quality CI gates.


Cloud-Native Infrastructure, Containers & SRE Operations
Cloud-Native Infrastructure, Kubernetes & DevOps Validation
Architected and hardened production-grade container infrastructure: built multi-stage Docker images (vLLM/PyTorch inference optimization), Kubernetes orchestration (Sidecar logging, VolumeMounts, HPA auto-scaling), zero-trust VPC peering, IaC Terraform/CloudFormation, FinOps cost governance, and Prometheus/Grafana telemetry.
- Containers & K8s: Multi-stage Docker builds (vLLM/PyTorch inference optimization), Kubernetes orchestration (Sidecar logging, VolumeMounts, HPA scaling).
- IaC & Security: Zero-trust VPC peering, IaC Terraform/CloudFormation, FinOps cost optimization, and HashiCorp Vault security.
- Observability: Strict Prometheus and Grafana platform observability.


Zero-Day Vulnerability & Advanced Penetration Research
Offensive Security Research (Independent)
Executed technical audits and independent threat-modeling protocols targeting cryptography, vulnerability exploitation, and network security.
- Defensive Frameworks: Applied defensive protocols (Burp Suite / Wireshark) to validate Zero-Trust data streams.
- Threat Modeling: Validated security posture against production-grade infiltration tactics.


Large Language Models & RAG Architectures
LangChain & Vector Databases for LLM Applications
Specialized training in building/deploying advanced LLM applications, RAG systems, and agentic workflows.
- Vector Search: Applied vector retrieval optimization using Pinecone.
- Memory Persistence: Implemented DeepLake memory persistence protocols to establish forensic state analysis for LLMs.

Contact

Available for enterprise-scale AI/ML architecture projects, fractional CTO roles, and applied MLOps consulting.Availability & Location Alignment
USA (EST) Overlap
Full European (CET/GMT) Alignment
Get in touch!