Lead Site Reliability Engineer
Aug 2025 – PresentViable Solutions (via Synechron, for Latitude Financial Services) | Permanent
- Team leadership: Lead a team of 10 offshore engineers in India, bridging client expectations and offshore delivery for infrastructure support across 80+ AWS accounts.
- Support model: Streamlined intake for all additional workloads through ServiceNow and set up a follow-the-sun model across Melbourne and India for 24x7 support.
- AI SRE (Alfred): Designed and built a multi-agent incident investigator, now live in production on Amazon Bedrock AgentCore, orchestrating AWS DevOps Agent with Datadog, Dynatrace, Buildkite and GitHub, with dependency-graph blast radius on pull requests.
- Claude for developers: Building an OpenAI-compatible Lambda proxy over Bedrock so GitHub Copilot (BYOK) can use Claude models, with a mandatory server-side guardrail and per-developer keys for cost attribution.
- ECS upgrade automation: Automated ECS cluster upgrades across all environments with zero manual intervention, moving clusters to CIS-hardened Amazon Linux 2023 AMIs.
- Patch automation: Removed ClickOps from monthly patching by introducing Ivanti Security Controls; designed the networking and architecture the team built on.
- Observability consolidation: Migrated 20M+ log events per week from Sumo Logic to Datadog, including Grok parsing pipelines for payment logs and transaction-orphan detection monitors, for single-pane dashboards and lower MTTR.
- Security: Rolled out CrowdStrike Falcon sensor injection for ECS Fargate workloads using the init-container model.
- High availability: Designed a 2-node active/passive Windows Failover Cluster on shared EBS io2 Multi-Attach to replace single-server Control-M staging hosts across Test, Pre-Prod and Prod.
- IaC adoption: Moving AWS accounts inherited from a previous vendor from ClickOps to CloudFormation and Terraform.