# Alfred: a multi-agent AI SRE

Latitude Financial Services (client, via Viable Solutions and Synechron) · Shipped · Role: Designed and built it, from prototype to production

**Live in production on Amazon Bedrock AgentCore, orchestrating AWS DevOps Agent**

A multi-agent incident investigator that works across AWS, Datadog, Dynatrace, Buildkite and GitHub in parallel and correlates the findings into one timeline and root cause. It started as a local helper; I took it to production on Amazon Bedrock AgentCore.

## Key facts

- **Where**: Latitude Financial Services (client), as Lead SRE at Viable Solutions via Synechron
- **Status**: Live in production since September 2026
- **Platform**: Amazon Bedrock AgentCore, orchestrating AWS DevOps Agent

## Problem

Diagnosing an incident meant stitching together several platforms by hand while it was still burning. That skill sat with a few engineers, and what they worked out was rarely written down, so the same failures were solved again later.

## Architecture

Production architecture, simplified. Client systems, data and access controls are deliberately not shown.

1. **Entry points**: Slack / Teams (Ask during an incident); Alerts (Investigations start automatically); IDE (Deep investigation); Pull requests (Blast-radius comments)
2. **Orchestrate**: Alfred orchestrator (AgentCore Runtime; Bedrock inference in-account)
3. **Investigate in parallel**: AWS DevOps Agent (AWS-side investigation); Specialist agents (Datadog, Dynatrace, Buildkite, GitHub via AgentCore Gateway); Dependency graph (Deterministic blast radius)
4. **Answer and learn**: Timeline and root cause (Findings correlated, blind spots declared); Knowledge base (Updated after every investigation)

## How it evolved

### Prototype: a local multi-agent helper (March 2026)

- An orchestrator in the IDE dispatching specialist agents for AWS, Datadog, Buildkite and GitHub over MCP.
- A structured investigation loop: recall similar past incidents, triage, form 3 to 5 hypotheses, fan out one hypothesis per agent, correlate, confirm, diagnose, then record what was learned.
- Used on real production investigations, which proved the pattern but kept the value on one machine.

### Production: hosted on Amazon Bedrock AgentCore (September 2026)

- Agents run on AgentCore Runtime, with tools exposed through AgentCore Gateway and AgentCore Memory and Identity services.
- Alfred orchestrates AWS DevOps Agent as its AWS specialist, alongside agents for Datadog, Dynatrace, Buildkite and GitHub.
- Reachable where engineers already work: Slack and Teams, the IDE, pull requests, and automatically from alerts.
- A dependency graph adds blast-radius comments to pull requests.

## My contribution

- Designed and built the original multi-agent prototype: orchestrator, specialist agents, MCP integrations and the knowledge-base loop.
- Productionised it on Amazon Bedrock AgentCore, with Alfred orchestrating AWS DevOps Agent.
- Built the dependency graph behind pull-request blast-radius comments.

## Key decisions

- An orchestrator with one specialist agent per platform, run in parallel, one hypothesis per dispatch.
- AWS DevOps Agent as the AWS specialist, orchestrated alongside Alfred’s own agents.
- Inference through Amazon Bedrock inside the company AWS account.

## Trade-offs

- MCP over CLI tools: one setup and read-only enforcement on the server side, instead of installing and authenticating four CLIs and trusting the prompt. Responses are structured, which also keeps agent context small.
- Graph, not model, for blast radius: a deterministic dependency graph decides what a change reaches, so results are citable and cheap. The model only maps a diff onto the graph, ranks risk and explains it.
- Declared blind spots over confident answers: every report states what it could not see. That reads as less certain, but a tool that misses something while sounding sure teaches people to stop checking.

## Implementation

- Orchestrator and specialist agents on AgentCore Runtime; tools through AgentCore Gateway.
- Integrations with AWS (via AWS DevOps Agent), Datadog, Dynatrace, Buildkite and GitHub.
- Dependency graph for pull-request blast radius.

## Reliability and safety

- Read-only, enforced in three layers: agent instructions, connector and gateway configuration, and IAM.
- Inference stays inside the company AWS account.
- Every investigation enriches a reviewed knowledge base, so fixes are not rediscovered.

## Outcomes

- Live in production since September 2026.
- Used from Slack and Teams, the IDE, pull requests, and automatically from alerts.
- Blast-radius comments on pull requests from the dependency graph.

## Stack

Amazon Bedrock AgentCore, AWS DevOps Agent, MCP, Datadog, Dynatrace, Buildkite, GitHub

_The production code is internal to the client. This page describes public AWS services and design principles only; no client data, systems or incidents are shown._

Web page: https://naman-kumar2397.github.io/projects/alfred/ · Author: Naman Kumar, Lead Site Reliability Engineer (naman.kumar2397@gmail.com)
