Laxmi Narasimha Hari Yelesetty Lead AI Engineer  ·  LLM Systems & Production GenAI  ·  GCP · Azure · GPU Infra
ylnharimailme@gmail.com
+91 99161 12881
linkedin.com/in/ylnhari
github.com/ylnhari
ylnhari.github.io
LLM & GenAI
vLLMUnsloth LoRA / QLoRAPEFT Transformers HF Accelerate Self-hosted LLM Ops RAG Prompt Engineering LLM Evaluation
ML & Deep Learning
PyTorchTensorFlow KerasScikit-learn RAPIDS (cuDF/CuPy) CUDA NVIDIA Morpheus Computer Vision NLPTime-Series Anomaly Detection
MLOps & Platforms
Cloud & Orchestration
GCP Vertex AI GKEGCS Azure ML Studio Kubeflow Azure DevOps
Serving & Packaging
KubernetesHelm IstioDocker FastAPIMLflow Artifact Registry
Observability & CI/CD
GrafanaPrometheus GMP PodMonitoring Cloud Build GitHub Actions
Languages & Data
PythonSQL PandasNumPy MS SQL ServerIBM DB2 KQL (ADX)
Certifications
Oracle CloudGenAI Certified Professional
Microsoft CertifiedAzure AI Engineer Associate
Microsoft CertifiedAzure Data Scientist Associate
Microsoft CertifiedPower BI Data Analyst Associate
Microsoft CertifiedPower Platform Fundamentals
Coursera / deeplearning.aiNeural Networks & Deep Learning
Education
B.E. Electrical & Electronics Andhra University, Visakhapatnam 2012 – 2016  ·  80%

A production AI/ML systems engineer with 10+ years taking models from the research notebook to live customer traffic. At Best Buy I built a config-driven pipeline that fine-tunes open-source language models (LoRA) and serves them in-house — replacing paid external AI APIs to save $300K/year — powering real-time fraud detection and personalization at 40,000+ transactions/sec. Earlier, at Wipro, I built GPU-accelerated cybersecurity ML on an NVIDIA DGX supercomputer (8×A100): models that learn each cloud account's normal behaviour from billions of rows of security logs to surface intrusions; a phishing-email detector running at 2× the previous speed; and automated triage of the customer "abuse" mailbox — where shoppers forward suspected scams and fraud — handling 2× the case volume and routing each threat to the right security team. Before that, an Azure ML system predicting product quality on arrival saved $2.5M/year, and Google OR-Tools route optimization streamlined last-mile delivery. Across 10 years and six companies I've earned the top performance rating at every one — the specialist for the gap between works in a notebook and serves real traffic.

Work Experience
Lead Artificial Intelligence Engineer
Best Buy  ·  Bengaluru, India
Jul 2025 – Present
  • Built config-driven LoRA fine-tuning framework — Unsloth + PEFT dual-backend (~2× faster, 60–80% less VRAM via QLoRA), single/multi-GPU via HF Accelerate — enabling zero-Python task addition via YAML config and multi-task single-adapter training.
  • Deployed vLLM OpenAI-compatible serving on GKE with HPA autoscaling and Istio traffic routing; GCS model store, Artifact Registry versioning, Cloud Build CI, GMP PodMonitoring/Prometheus/Grafana observability layer.
  • Self-hosted open-source LLMs end-to-end, eliminating external API dependency and saving $300K while strengthening data governance and compliance posture.
  • Productionalizing GenAI solutions on GCP (Vertex AI) for search, personalization, fraud detection, and cybersecurity at 40k+ transactions/sec; defining AI engineering standards and deployment best practices for retail-scale inference.
ML Lead Engineer — ML Engineering & Operations
Wipro  ·  Client: Major US Retail Corporation
Aug 2023 – Jul 2025
  • Built behavioral anomaly detection for AWS account security: trained unsupervised autoencoder models per entity (account / user / service) on NVIDIA DGX (8×A100) via NVIDIA Morpheus — learning normal activity fingerprints and surfacing deviations in real time. 20+ days of logs across multiple AWS accounts = billion-row datasets per training run.
  • Built phishing detection (NVIDIA Morpheus) scoring 2,400 emails/sec — 2× the prior throughput; built automated abuse-inbox triage handling 2× the prior pipeline's case volume — classifying customer-forwarded suspicious mail (OCR-extracted scammer identifiers) and auto-routing each threat to the right security response team.
  • Designed RAPIDS (cuDF/CuPy) GPU data engineering pipelines across all cybersecurity workloads; orchestrated batch training and inference on the DGX cluster; anomaly alerts surfaced via Grafana.
Application Dev Team Lead — Data Science / Azure ML
Accenture
May 2021 – Aug 2023
  • Developed end-to-end Azure ML pipeline for warehouse product quality prediction (BamaGruppen, Norway) — $2.5M annual savings, trucks inspected/day 130 → 160, +17% defect detection, +20% customer approval; stack: Azure ML Studio, Logic Apps, Azure Pipelines, Azure Storage.
  • Led last-mile delivery optimization using OR-Tools (Vehicle Routing Problem) — improving route planning, departure scheduling, and load distribution with significant annual logistics cost savings.
  • Designed cross-functional supply chain simulation framework for strategic hub location optimization, improving distribution efficiency and reducing the global carbon footprint across the client's network.
Earlier Experience
Data Scientist GD Research Centre Pvt Ltd May 2020 – May 2021 Built company-wide time-series forecasting across multiple intelligence centers (Consumer Goods, Oil & Gas, Power, Retail, Finance).
Deep Learning Engineer Tata Consultancy Services Jul 2019 – Mar 2020 Built & deployed hybrid cloud + on-prem CV system for multi-class semiconductor defect classification (DenseNet, transfer learning) at Amkor's South Korea site — eliminated manual inspection bottlenecks.
Senior Software Engineer — ML Infosys May 2016 – Jul 2019 Built ML-powered Lead Conversion Prediction system and Business Intelligence applications using Microsoft BI tools.
Selected Projects
LLM Fine-tuning & Serving Framework Best Buy  ·  Jul 2025 – Present

Config-driven LoRA adapter training for open-source causal LLMs with YAML-driven task addition — no Python required for new use cases. Unsloth backend for ~2× faster training and 60-80% VRAM reduction via QLoRA; HF Accelerate for single/multi-GPU. vLLM serving on GKE with hot-swap adapter support, HPA autoscaling, Istio routing, and GCS-backed model store.

UnslothPEFTvLLM GKEHF AccelerateGCS
$300K saved  ·  ~2× training speedup  ·  60-80% VRAM reduction
AWS Account Behavioral Fingerprinting Wipro  ·  Aug 2023 – Jul 2025

Unsupervised autoencoder models trained per-entity (account / user / service / machine) to profile normal AWS activity patterns and surface anomalous deviations in real time. Training and inference pipelines engineered with NVIDIA Morpheus on DGX (8×A100). 20+ days of logs across multiple AWS accounts totaling billion-row datasets per run; model management via MLflow; alerts served via Grafana.

NVIDIA MorpheusAutoencoders RAPIDS / cuDFMLflowGrafana
Billion-row datasets/run  ·  Per-entity autoencoders  ·  8×A100 DGX
Last-Mile Delivery Optimization Accenture  ·  May 2021 – Aug 2023

Vehicle Routing Problem (VRP) formulation using Google OR-Tools for a major logistics operator — optimizing route planning, departure scheduling, and product load distribution across a delivery fleet. Constraint modelling accounted for time windows, vehicle capacity, and geographic clustering.

OR-ToolsVRP PythonSimulation
Significant annual logistics cost savings  ·  OR-Tools VRP at fleet scale
Product Quality Prediction — BamaGruppen Accenture (Norway)  ·  May 2021 – Aug 2023

End-to-end Azure ML pipeline predicting product quality at warehouse arrival for a Norwegian retail group — enabling proactive inspection scheduling and reducing downstream quality failures. Fully operationalized with automated triggers and pipeline orchestration.

Azure ML StudioLogic Apps Azure PipelinesAzure Storage
$2.5M saved  ·  130 → 160 inspected/day  ·  +17% detection  ·  +20% approval