naman-kumar / journey / projects / alfredPublic

Back to all projects

ShippedLatitude Financial Services (client, via Viable Solutions and Synechron)

Alfred: a multi-agent AI SRE

Live in production on Amazon Bedrock AgentCore, orchestrating AWS DevOps Agent

A multi-agent incident investigator that works across AWS, Datadog, Dynatrace, Buildkite and GitHub in parallel and correlates the findings into one timeline and root cause. It started as a local helper; I took it to production on Amazon Bedrock AgentCore.

My role
Designed and built it, from prototype to production
Where
Latitude Financial Services (client), as Lead SRE at Viable Solutions via Synechron
Status
Live in production since September 2026
Platform
Amazon Bedrock AgentCore, orchestrating AWS DevOps Agent

Architecture

  1. Entry points

    • Slack / TeamsAsk during an incident
    • AlertsInvestigations start automatically
    • IDEDeep investigation
    • Pull requestsBlast-radius comments
  2. Orchestrate

    • Alfred orchestratorAgentCore Runtime; Bedrock inference in-account
  3. Investigate in parallel

    • AWS DevOps AgentAWS-side investigation
    • Specialist agentsDatadog, Dynatrace, Buildkite, GitHub via AgentCore Gateway
    • Dependency graphDeterministic blast radius
  4. Answer and learn

    • Timeline and root causeFindings correlated, blind spots declared
    • Knowledge baseUpdated after every investigation
Production architecture, simplified. Client systems, data and access controls are deliberately not shown.

How it evolved

  1. Prototype: a local multi-agent helper · March 2026

    • An orchestrator in the IDE dispatching specialist agents for AWS, Datadog, Buildkite and GitHub over MCP.
    • A structured investigation loop: recall similar past incidents, triage, form 3 to 5 hypotheses, fan out one hypothesis per agent, correlate, confirm, diagnose, then record what was learned.
    • Used on real production investigations, which proved the pattern but kept the value on one machine.
  2. Production: hosted on Amazon Bedrock AgentCore · September 2026

    • Agents run on AgentCore Runtime, with tools exposed through AgentCore Gateway and AgentCore Memory and Identity services.
    • Alfred orchestrates AWS DevOps Agent as its AWS specialist, alongside agents for Datadog, Dynatrace, Buildkite and GitHub.
    • Reachable where engineers already work: Slack and Teams, the IDE, pull requests, and automatically from alerts.
    • A dependency graph adds blast-radius comments to pull requests.

Problem

Diagnosing an incident meant stitching together several platforms by hand while it was still burning. That skill sat with a few engineers, and what they worked out was rarely written down, so the same failures were solved again later.

My contribution

  • Designed and built the original multi-agent prototype: orchestrator, specialist agents, MCP integrations and the knowledge-base loop.
  • Productionised it on Amazon Bedrock AgentCore, with Alfred orchestrating AWS DevOps Agent.
  • Built the dependency graph behind pull-request blast-radius comments.

Engineering decisions

  • An orchestrator with one specialist agent per platform, run in parallel, one hypothesis per dispatch.
  • AWS DevOps Agent as the AWS specialist, orchestrated alongside Alfred’s own agents.
  • Inference through Amazon Bedrock inside the company AWS account.

Trade-offs

  • MCP over CLI tools: one setup and read-only enforcement on the server side, instead of installing and authenticating four CLIs and trusting the prompt. Responses are structured, which also keeps agent context small.
  • Graph, not model, for blast radius: a deterministic dependency graph decides what a change reaches, so results are citable and cheap. The model only maps a diff onto the graph, ranks risk and explains it.
  • Declared blind spots over confident answers: every report states what it could not see. That reads as less certain, but a tool that misses something while sounding sure teaches people to stop checking.

Reliability and security

  • Read-only, enforced in three layers: agent instructions, connector and gateway configuration, and IAM.
  • Inference stays inside the company AWS account.
  • Every investigation enriches a reviewed knowledge base, so fixes are not rediscovered.

Verified outcomes

  • Live in production since September 2026.
  • Used from Slack and Teams, the IDE, pull requests, and automatically from alerts.
  • Blast-radius comments on pull requests from the dependency graph.

Stack

The production code is internal to the client. This page describes public AWS services and design principles only; no client data, systems or incidents are shown.

git log project/alfred: see this in the career journey