Amit KumarAmit Kumar

Work

What I built, what it was for, and what changed because of it.

  • Agentic FinOps intelligence layer for Azure, in GoA two-stage LLM pipeline: an anomaly-detection agent correlates spend spikes with infrastructure change events, and a synthesizer explains root cause in plain English.AgenticSHIPPED
  • The Kubernetes foundation for a multi-cloud healthcare platformBuilt the AKS cluster on Azure + AWS as the shared foundation applications and internal tools run on — so teams get a consistent place to deploy instead of one-off setups.Evolent2024 — NOW
  • A self-service developer portal with BackstageHelped bring Backstage to the org — a self-service portal so engineers can ship on their own instead of waiting on the platform team, backed by GitOps delivery and reusable release workflows.PlatformSHIPPED
  • Full-stack LGTM observability with OpenTelemetryLoki, Grafana, Tempo, and Mimir on AKS with OTel and Pydantic LogFire for AI-ready ops pipelines. Same visibility, a fraction of the managed bill.−60%SPEND
  • AIOps foundation: KubeFlow pipelines and KServe on AKSDeployed ML pipelines and KServe inference (Gemma) as the base for automated incident prediction and anomaly detection, aligned with SRE practice — SLOs, incident response, post-mortems.AIOpsSHIPPED
  • A Go REST API over Terragrunt and TerraformProgrammatic infrastructure provisioning via API calls, integrated with internal platform pipelines instead of CLI-based workflows.IaCAPI
  • Azure OpenAI and Document Intelligence via TerraformDeployed for business-ops analytics and log anomaly detection, with custom content filters configured for compliance.AISHIPPED
  • Ephemeral ARC runners on KubernetesGitHub Actions Runner Controller for self-hosted runners as throwaway pods; cut CI/CD compute spend without slowing delivery.−30%SPEND
  • Kyverno image signing and supply-chain integrityHelped build the container security baseline — admission policy and image signing across AKS, so only trusted images reach production.Supply chainHARDENED
  • ELK at 60 GiB/day and reusable Terraform modulesRan three ELK clusters processing 60 GiB/day for an e-commerce product across 8+ European countries, and centralized the Terraform modules nobody wanted to hand-roll twice.Telekom2022 — 24
  • Helm-templatized 40+ microservices across four EKS clustersOne template across all environments, deployed with zero issues. Central shared Jenkins, Kafka, and Redis as platform services.ScaleZERO-ISSUE
  • RHEL 6→8 migration across a 400+ server estateProvisioned large-scale highly-available systems on AWS and on-prem for Lifetouch; automated vulnerability management and patching with Ansible.ValueLabs2020 — 22