Naman Kumar

Lead Site Reliability Engineer · Melbourne, VIC

naman.kumar2397@gmail.com · linkedin.com/in/-namankumar · naman-kumar2397.github.io

Summary

Lead Site Reliability Engineer with 7+ years of hands-on experience designing, scaling and automating large-scale distributed systems on AWS. Currently leading a team of 10 engineers supporting 80+ AWS accounts for a major financial services client. Proven track record in cutting incident response times, consolidating observability platforms, and moving ClickOps estates to Infrastructure as Code. Now focused on AI-driven operations, including LLM-powered incident tooling and secure enterprise adoption of Claude on AWS Bedrock.

Skills

Work Experience

Lead Site Reliability Engineer

Aug 2025 – Present

Viable Solutions (via Synechron, for Latitude Financial Services) | Permanent

  • Team leadership: Lead a team of 10 offshore engineers in India, bridging client expectations and offshore delivery for infrastructure support across 80+ AWS accounts.
  • Support model: Streamlined intake for all additional workloads through ServiceNow and set up a follow-the-sun model across Melbourne and India for 24x7 support.
  • AI SRE (Alfred): Designed and built a multi-agent incident investigator, now live in production on Amazon Bedrock AgentCore, orchestrating AWS DevOps Agent with Datadog, Dynatrace, Buildkite and GitHub, with dependency-graph blast radius on pull requests.
  • Claude for developers: Building an OpenAI-compatible Lambda proxy over Bedrock so GitHub Copilot (BYOK) can use Claude models, with a mandatory server-side guardrail and per-developer keys for cost attribution.
  • ECS upgrade automation: Automated ECS cluster upgrades across all environments with zero manual intervention, moving clusters to CIS-hardened Amazon Linux 2023 AMIs.
  • Patch automation: Removed ClickOps from monthly patching by introducing Ivanti Security Controls; designed the networking and architecture the team built on.
  • Observability consolidation: Migrated 20M+ log events per week from Sumo Logic to Datadog, including Grok parsing pipelines for payment logs and transaction-orphan detection monitors, for single-pane dashboards and lower MTTR.
  • Security: Rolled out CrowdStrike Falcon sensor injection for ECS Fargate workloads using the init-container model.
  • High availability: Designed a 2-node active/passive Windows Failover Cluster on shared EBS io2 Multi-Attach to replace single-server Control-M staging hosts across Test, Pre-Prod and Prod.
  • IaC adoption: Moving AWS accounts inherited from a previous vendor from ClickOps to CloudFormation and Terraform.

DevOps Lead Consultant

May 2025 – Jun 2025

CI&T (for iSelect / CompareTheMarket) | Contract

  • Merger: Consolidated technology assets during the merger of iSelect and CompareTheMarket.
  • IaC migration: Led the move from CloudFormation to Terraform across multiple AWS accounts with shared modules and state management.
  • Data platform: Integrated ingestion pipelines from Salesforce, Meta, Google Ads and third-party SFTP APIs into a unified Delta Lakehouse on Databricks, using Python, Lambda, DMS and AWS Transfer Family.

Lead Site Reliability Engineer

Sep 2024 – Apr 2025

AKQA

  • Netflix game launch: Led SRE for an AI-driven face-transforming game, designing Kubernetes clusters that scaled to 50k+ concurrent users in the first 10 minutes of launch, with Grafana monitoring of NVIDIA DCGM GPU metrics.
  • Observability revamp: Reviewed and upgraded monitoring, logging and alerting across all client infrastructures.
  • Incident automation: Built an incident management workflow in Python and AWS Lambda integrating OpsGenie, Slack, Teams and Jira, reducing manual overhead and response times.
  • Change management: Defined Change Request standards and automated approvals to reduce deployment risk.

Senior Site Reliability Engineer

Oct 2019 – Sep 2024

Cvent

  • Incident automation: Co-developed Sev1 incident response automation (Datadog and Slack triggers orchestrating Jira, Slack, Zoom and PagerDuty), cutting time-to-respond by 20+ minutes.
  • Log migration: Migrated ~880M log events per week from Splunk to Datadog, unifying logs, metrics and traces on one platform.
  • Observability ownership: Owned SLI/SLO and custom instrumentation standards for Java microservices; led a newly acquired company's move to Datadog (10+ dashboards).
  • Platform scale: Rolled out ECS Capacity Providers across 70+ clusters with zero downtime.
  • Octopus Deploy: Built a high-availability model sustaining 99.9% availability and cut self-hosting costs by 30%.
  • Cloud migration: Contributed to migrating a Rackspace-hosted product (17 Java microservices, 26 .NET projects, 8 websites) to AWS, and standardised build pipelines for 134 .NET projects across Cvent onto one pipeline.
  • IaC: Early adopter of AWS CDK, consolidating CloudFormation and Terraform stacks into a single CDK stack for multiple applications.
  • Security: Moved all microservices to non-root, least-privilege containers.
  • Developer productivity: Embedded with sprint teams to remove bottlenecks, including fixing flaky JUnit/Jest tests that were wasting PR build hours; ran Fargate and Harness POCs.
  • Configuration management: Re-engineered Chef for 500+ monthly active servers.
  • Mentoring: Onboarded and mentored 4 new engineers.

Cloud Automation Associate

Jan 2019 – Oct 2019

Infraguard

  • Wrote Lambda automations for scheduled start/stop, package updates and SSH key rotation.
  • Designed a NAT Gateway pattern to patch sensitive private-subnet servers without exposing them to the internet.
  • Built the knowledge base for multi-cloud VM management across AWS, GCP, Azure and Alibaba Cloud.

Projects

Accolades (Cvent)

Education

Bachelor of Technology, Manipal University Jaipur (2015 – 2019)

Certifications

AWS Certified Solutions Architect – Associate (R2XYHRM1JE1E1TKJ), earned 2019