Enterprise platform scale
Designing self-service tooling and shared platform capabilities for 60+ managed services.

I build Kubernetes, CI/CD, and observability platforms that help engineering teams ship faster and operate reliably — supporting 60+ managed services across multiple regions and reducing incident detection time by 70%.
Currently deep in Kubernetes
I work across the complete delivery lifecycle: developer enablement, infrastructure automation, Kubernetes operations, observability, incident response, and continuous reliability improvement.
Designing self-service tooling and shared platform capabilities for 60+ managed services.
Removing manual work with CI/CD, Python, Bash, Terraform, and Ansible automation.
Building actionable Prometheus and Grafana observability for multi-region Kubernetes.
Leading incident investigation, root-cause analysis, resilience testing, and recovery.
I'm a Senior DevOps & Platform Engineer based in Pune, India, with 7+ years of experience in CI/CD, infrastructure automation, and observability. I work across the full delivery lifecycle — from building and packaging apps with Docker, to shipping them through CI/CD, to running and monitoring them on multi-region Kubernetes.
I enjoy automation and platform tooling: writing Python and Bash to remove toil, authoring infrastructure-as-code with Terraform and Ansible, building developer-facing services with FastAPI, and designing Prometheus/Grafana observability for production systems.
I've completed a Post Graduate Program in Cloud Computing and enjoy writing hands-on DevOps learning guides to share what I learn.
The tools and technologies I use to build, ship, and run software.
7+ years across platform engineering, DevOps, and IT operations.
Amdocs
Jan 2022 – Present
Cognizant
Nov 2019 – Jan 2022
Validated cloud expertise and formal training.
Open-source projects I've built and shared on GitHub — plus a deep dive into platform work delivered at enterprise scale.
A single pane of glass for a platform team — real-time Kubernetes/OpenShift pod health, MCP microservice usage analytics, support-ticket trends, and user adoption behind one FastAPI backend, with rate-aware health alerting and a bundled demo mode that runs with zero credentials.
A 12-lab, clone-and-go practice collection spanning Linux, Python, Git, Docker, Kubernetes, Terraform, Ansible, CI/CD, Cloud, and System Design — each with hands-on exercises, solutions, cheatsheets, theme-aware diagrams, and interview Q&A.
A 100% local, offline AI tutor for DevOps that turns your own practice labs into an interactive coach — adaptive study, mock interviews, and SM-2 spaced repetition. Built with a local RAG pipeline (Ollama + Chroma) so nothing leaves your machine — no API keys, no cloud bills.
Consolidated fragmented, region-by-region monitoring and manual operations into a single self-service platform for 60+ managed services across multiple regions — giving teams unified visibility and cutting incident detection time by 70%.
Fragmented monitoring across regions meant slow, inconsistent incident detection and heavy manual operations.
Centralized Prometheus/Grafana with multi-region alerting on replication, sync health, and capacity — fronted by self-service tooling on Kubernetes.
60+ services covered, incident detection 70% faster, and repeatable provisioning via 17+ automated pipelines.
Hands-on DevOps guides I've written to teach by doing — and to consolidate what I learn.
From first container to multi-service orchestration, with practical labs and gotchas.
Declarative pipelines, shared libraries, and patterns for reliable delivery.
Repeatable, multi-environment infrastructure and configuration management.
Metrics, dashboards, and alerting that actually catch problems early.
Open to DevOps, Cloud, Platform, and SRE opportunities. Have a project or a role in mind? Let's talk.