Shubham Sharma
Senior Platform / DevOps Engineer — Pune, India

Shubham Sharma
Platform engineering at enterprise scale.

I build Kubernetes, CI/CD, and observability platforms that help engineering teams ship faster and operate reliably — supporting 60+ managed services across multiple regions and reducing incident detection time by 70%.

Currently deep in Kubernetes

Open to Senior Platform / DevOps opportunities
What I bring

Platform thinking backed by production ownership.

I work across the complete delivery lifecycle: developer enablement, infrastructure automation, Kubernetes operations, observability, incident response, and continuous reliability improvement.

01

Enterprise platform scale

Designing self-service tooling and shared platform capabilities for 60+ managed services.

02

Delivery automation

Removing manual work with CI/CD, Python, Bash, Terraform, and Ansible automation.

03

Reliability by design

Building actionable Prometheus and Grafana observability for multi-region Kubernetes.

04

Production ownership

Leading incident investigation, root-cause analysis, resilience testing, and recovery.

01 — About

Who I am

I'm a Senior DevOps & Platform Engineer based in Pune, India, with 7+ years of experience in CI/CD, infrastructure automation, and observability. I work across the full delivery lifecycle — from building and packaging apps with Docker, to shipping them through CI/CD, to running and monitoring them on multi-region Kubernetes.

I enjoy automation and platform tooling: writing Python and Bash to remove toil, authoring infrastructure-as-code with Terraform and Ansible, building developer-facing services with FastAPI, and designing Prometheus/Grafana observability for production systems.

I've completed a Post Graduate Program in Cloud Computing and enjoy writing hands-on DevOps learning guides to share what I learn.

7+
Years experience
60+
Managed services supported
17+
CI/CD pipelines automated
40%
Fewer vulnerabilities
02 — Skills

What I work with

The tools and technologies I use to build, ship, and run software.

Containers & Orchestration

DockerKubernetesHelm

CI/CD & IaC

JenkinsAzure DevOps TerraformAnsibleSonarQube

Languages & Scripting

PythonBash GroovyYAMLLinux

Observability

PrometheusGrafana DatadogSplunkNginx

Cloud

AWSAzure

Artifact & SCM

NexusArtifactory GitBitbucketPerforce
03 — Experience

Where I've worked

7+ years across platform engineering, DevOps, and IT operations.

Amdocs Jan 2022 – Present
Pune, India
Senior DevOps Engineer Jun 2024 – Present
  • Designed enterprise DevOps enablement & observability platforms supporting 60+ managed services, improving operational visibility and reducing incident detection time by 70%.
  • Built 17+ automated CI/CD pipelines in Python and Bash for infrastructure provisioning, compliance verification, and deployment workflows.
  • Implemented centralized multi-region monitoring and alerting with Prometheus and Grafana across Kubernetes platform services.
  • Led root-cause analysis, production incident troubleshooting, and disaster-recovery replication testing to sustain long-term infrastructure reliability.
KubernetesJenkinsPrometheusGrafanaAWSPython
DevOps Engineer Jan 2022 – Jun 2024
  • Built and maintained Jenkins and Azure DevOps pipelines supporting 20+ enterprise applications across development, staging, and production.
  • Containerized legacy applications with Docker and Kubernetes, scaling deployment infrastructure through Ansible playbook automation.
  • Administered SonarQube quality gates and Nexus repositories, introducing governance controls that reduced reported vulnerabilities by 40%.
Azure DevOpsDockerAnsibleSonarQubeNexus
Cognizant Nov 2019 – Jan 2022
Pune, India
Associate DevOps Engineer Jun 2021 – Jan 2022
  • Built Jenkins CI/CD workflows and optimized release management, reducing manual deployment effort by 60%.
  • Integrated Splunk monitoring and ServiceNow ticket routing to accelerate operational visibility and ensure strict SLA compliance.
JenkinsSplunkServiceNow
Programmer Analyst — IT Operations Nov 2019 – Jun 2021
  • Managed critical production incidents with structured root-cause analysis and risk-mitigated change management using ITIL-aligned processes.
  • Developed custom Python and Bash automation to minimize operational toil, improving incident resolution times by 40%.
PythonBashITIL
04 — Credentials

Certifications & education

Validated cloud expertise and formal training.

Education
Post Graduate Program in Cloud Computing
Great Learning · 2025
B.E., Electronics & Communication
RGPV, Bhopal · 2014 – 2018
05 — Selected Work

Selected work

Open-source projects I've built and shared on GitHub — plus a deep dive into platform work delivered at enterprise scale.

PLATFORM · OBSERVABILITY

Operations Dashboard

A single pane of glass for a platform team — real-time Kubernetes/OpenShift pod health, MCP microservice usage analytics, support-ticket trends, and user adoption behind one FastAPI backend, with rate-aware health alerting and a bundled demo mode that runs with zero credentials.

FastAPIKubernetesChart.jsDockerObservability
EDUCATION · CONTENT

DevOps LearningHub

A 12-lab, clone-and-go practice collection spanning Linux, Python, Git, Docker, Kubernetes, Terraform, Ansible, CI/CD, Cloud, and System Design — each with hands-on exercises, solutions, cheatsheets, theme-aware diagrams, and interview Q&A.

KubernetesTerraformDockerCI/CDAnsible
AI · OPEN SOURCE

LabSensei

A 100% local, offline AI tutor for DevOps that turns your own practice labs into an interactive coach — adaptive study, mock interviews, and SM-2 spaced repetition. Built with a local RAG pipeline (Ollama + Chroma) so nothing leaves your machine — no API keys, no cloud bills.

PythonOllamaRAGChromaStreamlit
CASE STUDY

DevOps Enablement & Observability Platform

Consolidated fragmented, region-by-region monitoring and manual operations into a single self-service platform for 60+ managed services across multiple regions — giving teams unified visibility and cutting incident detection time by 70%.

Problem

Fragmented monitoring across regions meant slow, inconsistent incident detection and heavy manual operations.

Approach

Centralized Prometheus/Grafana with multi-region alerting on replication, sync health, and capacity — fronted by self-service tooling on Kubernetes.

Impact

60+ services covered, incident detection 70% faster, and repeatable provisioning via 17+ automated pipelines.

06 — Writing

Writing & guides

Hands-on DevOps guides I've written to teach by doing — and to consolidate what I learn.

CONTAINERS

Docker & Kubernetes

From first container to multi-service orchestration, with practical labs and gotchas.

CI/CD

Jenkins Pipelines

Declarative pipelines, shared libraries, and patterns for reliable delivery.

IAC

Terraform & Ansible

Repeatable, multi-environment infrastructure and configuration management.

OBSERVABILITY

Prometheus & Grafana

Metrics, dashboards, and alerting that actually catch problems early.

Let's build something reliable

Open to DevOps, Cloud, Platform, and SRE opportunities. Have a project or a role in mind? Let's talk.

— or reach me directly —