naman-kumar / journeyPublic

HEAD -> work/latitudeLead SRE at Viable Solutions, for client Latitude Financial Services

Naman Kumar

Lead Site Reliability Engineer Melbourne, VIC

Engineering reliable, scalable and observable systems at production scale.

I lead reliability engineering for large AWS estates: production operations, observability, infrastructure automation and the teams that run them. Lately that includes practical AI tooling for incident response on AWS Bedrock.

Portrait of Naman Kumar

Impact at a glance

git log --first-parent main

Experience

From cloud automation to leading an SRE team across 80+ AWS accounts. Consulting engagements show the employer, the consulting partner and the client separately.

Career at a glance

PeriodRoleEmployer and clientHighlight
Aug 2025 – PresentLead Site Reliability EngineerViable Solutions for Latitude Financial ServicesLeads a 10-engineer team supporting 80+ AWS accounts, 24x7
May 2025 – Jun 2025DevOps Lead ConsultantCI&T for iSelect / CompareTheMarketCloudFormation to Terraform through the iSelect / CompareTheMarket merger
Sep 2024 – Apr 2025Lead Site Reliability EngineerAKQASRE lead for a Netflix game launch: 50k+ concurrent users in 10 minutes
Oct 2019 – Sep 2024Senior Site Reliability EngineerCventCo-developed Sev1 automation (20+ min faster response); ~880M logs/week to Datadog
Jan 2019 – Oct 2019Cloud Automation AssociateInfraguardLambda automation and private-subnet patching across multi-cloud VMs
  1. Lead Site Reliability Engineer

    Current

    Aug 2025 – Present

    Employer
    Viable Solutions (permanent)
    Consulting partner
    Synechron
    Client
    Latitude Financial Services
    • Leads a team of 10 engineers
    • 80+ AWS accounts
    • 24x7 follow-the-sun support
    • Team leadership. Lead a team of 10 offshore engineers in India, bridging client expectations and offshore delivery for infrastructure support across 80+ AWS accounts.
    • AI SRE (Alfred). Designed and built a multi-agent incident investigator, now live in production on Amazon Bedrock AgentCore, orchestrating AWS DevOps Agent with Datadog, Dynatrace, Buildkite and GitHub, with dependency-graph blast radius on pull requests.
    • ECS upgrade automation. Automated ECS cluster upgrades across all environments with zero manual intervention, moving clusters to CIS-hardened Amazon Linux 2023 AMIs.
    More detail (7)
    • Observability consolidation. Migrated 20M+ log events per week from Sumo Logic to Datadog, including Grok parsing pipelines for payment logs and transaction-orphan detection monitors, for single-pane dashboards and lower MTTR.
    • Security. Rolled out CrowdStrike Falcon sensor injection for ECS Fargate workloads using the init-container model.
    • Support model. Streamlined intake for all additional workloads through ServiceNow and set up a follow-the-sun model across Melbourne and India for 24x7 support.
    • Claude for developers. Building an OpenAI-compatible Lambda proxy over Bedrock so GitHub Copilot (BYOK) can use Claude models, with a mandatory server-side guardrail and per-developer keys for cost attribution.
    • Patch automation. Removed ClickOps from monthly patching by introducing Ivanti Security Controls; designed the networking and architecture the team built on.
    • High availability. Designed a 2-node active/passive Windows Failover Cluster on shared EBS io2 Multi-Attach to replace single-server Control-M staging hosts across Test, Pre-Prod and Prod.
    • IaC adoption. Moving AWS accounts inherited from a previous vendor from ClickOps to CloudFormation and Terraform.

    Key platforms: AWS · ECS Fargate · Datadog · Dynatrace · ServiceNow · CrowdStrike · Terraform · Bedrock

  2. DevOps Lead Consultant

    May 2025 – Jun 2025

    Employer
    CI&T (contract)
    Client
    iSelect / CompareTheMarket
    • Multiple AWS accounts
    • Merger consolidation
    • Merger. Consolidated technology assets during the merger of iSelect and CompareTheMarket.
    • IaC migration. Led the move from CloudFormation to Terraform across multiple AWS accounts with shared modules and state management.
    • Data platform. Integrated ingestion pipelines from Salesforce, Meta, Google Ads and third-party SFTP APIs into a unified Delta Lakehouse on Databricks, using Python, Lambda, DMS and AWS Transfer Family.

    Key platforms: Terraform · Databricks · AWS DMS · AWS Transfer Family · Python

  3. Lead Site Reliability Engineer

    Sep 2024 – Apr 2025

    Employer
    AKQA
    • SRE lead for a Netflix game launch
    • All client infrastructures
    • Netflix game launch. Led SRE for an AI-driven face-transforming game, designing Kubernetes clusters that scaled to 50k+ concurrent users in the first 10 minutes of launch, with Grafana monitoring of NVIDIA DCGM GPU metrics.
    • Observability revamp. Reviewed and upgraded monitoring, logging and alerting across all client infrastructures.
    • Incident automation. Built an incident management workflow in Python and AWS Lambda integrating OpsGenie, Slack, Teams and Jira, reducing manual overhead and response times.
    More detail (1)
    • Change management. Defined Change Request standards and automated approvals to reduce deployment risk.

    Key platforms: Kubernetes · Grafana · NVIDIA DCGM · Python · AWS Lambda · OpsGenie

  4. Senior Site Reliability Engineer

    Oct 2019 – Sep 2024

    Employer
    Cvent
    • 70+ ECS clusters
    • 500+ monthly active servers (Chef)
    • Mentored 4 engineers
    • Incident automation. Co-developed Sev1 incident response automation (Datadog and Slack triggers orchestrating Jira, Slack, Zoom and PagerDuty), cutting time-to-respond by 20+ minutes.
    • Log migration. Migrated ~880M log events per week from Splunk to Datadog, unifying logs, metrics and traces on one platform.
    • Platform scale. Rolled out ECS Capacity Providers across 70+ clusters with zero downtime.
    More detail (8)
    • Octopus Deploy. Built a high-availability model sustaining 99.9% availability and cut self-hosting costs by 30%.
    • Cloud migration. Contributed to migrating a Rackspace-hosted product (17 Java microservices, 26 .NET projects, 8 websites) to AWS, and standardised build pipelines for 134 .NET projects across Cvent onto one pipeline.
    • Observability ownership. Owned SLI/SLO and custom instrumentation standards for Java microservices; led a newly acquired company's move to Datadog (10+ dashboards).
    • IaC. Early adopter of AWS CDK, consolidating CloudFormation and Terraform stacks into a single CDK stack for multiple applications.
    • Security. Moved all microservices to non-root, least-privilege containers.
    • Developer productivity. Embedded with sprint teams to remove bottlenecks, including fixing flaky JUnit/Jest tests that were wasting PR build hours; ran Fargate and Harness POCs.
    • Configuration management. Re-engineered Chef for 500+ monthly active servers.
    • Mentoring. Onboarded and mentored 4 new engineers.

    Key platforms: Datadog · Splunk · AWS ECS · AWS CDK · Octopus Deploy · PagerDuty · Chef

  5. Cloud Automation Associate

    Jan 2019 – Oct 2019

    Employer
    Infraguard
    • AWS, GCP, Azure and Alibaba Cloud
    • Wrote Lambda automations for scheduled start/stop, package updates and SSH key rotation.
    • Designed a NAT Gateway pattern to patch sensitive private-subnet servers without exposing them to the internet.
    • Built the knowledge base for multi-cloud VM management across AWS, GCP, Azure and Alibaba Cloud.

    Key platforms: AWS Lambda · VPC · Multi-cloud

ls projects/

Engineering case studies

Three flagship projects with full write-ups, plus current work in progress and a side project. Client-specific details are intentionally left out.

Side projects

  • Buildkite Build Watcher

    Open-source Chrome extension (Manifest V3), live on the Chrome Web Store. Plays a distinct synthesised chime when a build passes, fails or is blocked on input, and auto-watches builds you trigger.

    Independent project; not affiliated with Buildkite.

    • Chrome extension
    • Manifest V3
    • JavaScript
    • Buildkite

cat skills.yml

Core capabilities

In priority order, each backed by work described above.

  1. Site Reliability Engineering

    SLI/SLO standards, zero-downtime rollouts, Sev1 automation that cut response time by 20+ minutes

    • SLIs/SLOs
    • Incident response
    • High availability
    • PagerDuty
  2. Observability

    Two large log migrations to Datadog (~880M and 20M+ events per week); GPU monitoring for a launch

    • Datadog
    • Dynatrace
    • Grafana
    • Splunk
  3. Cloud Architecture on AWS

    80+ AWS accounts supported; 70+ ECS clusters; Rackspace to AWS migration

    • ECS / Fargate
    • Lambda
    • IAM
    • Kubernetes
  4. Infrastructure as Code and Platform

    ClickOps to IaC, CloudFormation to Terraform, Octopus Deploy HA at 99.9%

    • Terraform
    • AWS CDK
    • CloudFormation
    • CI/CD
  5. Technical Leadership

    Leads 10 engineers with a follow-the-sun 24x7 model; mentored 4 engineers

    • Team leadership
    • Support models
    • Change management
    • Mentoring
  6. AI-assisted Operations

    Alfred, a multi-agent AI SRE live on Bedrock AgentCore; a governed Claude proxy on Bedrock

    • Bedrock AgentCore
    • Claude
    • MCP
    • Guardrails
Full technology inventory (39 entries)
Cloud and Containers
AWS (ECS, Fargate, EC2, Lambda, S3, DMS, DynamoDB, IAM, KMS, Transit Gateway, Bedrock), Kubernetes, Docker
Infrastructure as Code
Terraform, AWS CDK, CloudFormation
CI/CD
Buildkite, GitLab, Jenkins, Harness, Octopus Deploy
Observability
Datadog, Dynatrace, Grafana, Sumo Logic, Splunk, New Relic, CloudWatch
AI and Automation
AWS Bedrock, Claude, MCP servers, Python automation
Security and Patching
CrowdStrike Falcon, Ivanti Security Controls, CIS-hardened AMIs, Least-privilege IAM
Incident and Service Management
PagerDuty, OpsGenie, ServiceNow, Jira
Configuration Management
Chef, Ansible
Languages
Python, TypeScript, Groovy, Shell, PowerShell
Version Control
GitHub, Bitbucket
  • AWS
  • Kubernetes
  • ECS
  • Docker
  • Fargate
  • Terraform
  • EC2
  • Datadog
  • Lambda
  • Dynatrace
  • S3
  • Grafana
  • DynamoDB
  • Splunk
  • DMS
  • Sumo Logic
  • IAM
  • New Relic
  • KMS
  • Claude
  • Transit Gateway
  • MCP
  • Bedrock
  • Buildkite
  • AWS CDK
  • GitLab
  • CloudFormation
  • Jenkins
  • CloudWatch
  • Octopus Deploy
  • Harness
  • PagerDuty
  • CrowdStrike
  • Opsgenie
  • Ivanti
  • Jira
  • ServiceNow
  • Chef
  • PowerShell
  • Ansible
  • Python
  • TypeScript
  • Groovy
  • Bash
  • GitHub
  • Bitbucket

git log --graph --all

Career journey

The full story, for the curious: my career as a git history. Each company is a branch off main that merged back when I moved on; projects branch off their company. Select any entry for details.

47 commits · 9 branches

  • Commit
  • Merge (moved on)
  • Award or certification
  • HEADWhere I am now

cat CONTACT.md

Get in touch

Email is the fastest way to reach me. I'm based in Melbourne, VIC.