Alexei Rojas Quiroga
← All projects

Corporate project · Presented company-agnostic: architecture and decisions only.

Agentic Software Delivery Platform

Persisted state machine coordinating AI agents from backlog item to approved pull request

Status
In production
Role
Lead engineer and architect: workflow state machine, LLM gateway, prompt-safety integration, MCP tool surface, vector memory, indexer design, portal and deployment shape.

Architecture

  • 4 role-specific AI agents
  • Persisted state machine with leases
  • Human approval gates
  • Work item to approved pull request

A ready signal creates a persisted run; the orchestrator dispatches definition, QA planning, plan approval, development, review and rework, then publishes a pull request, watches CI, and waits for human approval.

Architecture · 9 nodes · 2 flows
WHAT I BUILTAzure DevOpsWork tracking &reposItems, branches, PRsTeams botTeam chat botCommands, approvalcardsReact/SignalRAgent portalLive run view,approvalsCouchbaseRun & lease storeRuns, work, leases.NET 10WorkfloworchestratorPersisted state machineLLM agentReview agentAdversarial PR reviewLLM agentDefinition agentDrafts AI-ready itemsLLM agentQA agentTest plans and suitesLLM agentDev agentPlans, codes, tests1Work item ready signal(webhook)23456789
  • Service / compute
  • Data store
  • Client
  • External system
  • Synchronous
  • Async / loop
Scroll sideways to see the full diagram

How it flows, step by step

Click a step to jump to it. Click a component for details.

What it does

An internal platform that turns work-tracking items into reviewed pull requests by coordinating specialized AI agents for requirements, QA planning, development, adversarial review and CI repair. A multi-provider LLM gateway with a prompt-safety layer, an MCP server exposing scoped tools, vector memory and a code-to-graph indexer provide the shared infrastructure, with human approval gates at every critical step.

The problem

Teams wanted LLM-assisted delivery but could not allow agents to talk directly to providers, leak sensitive data in prompts, hold unrestricted repository or backlog access, or lose work when a long-running run crashed. The challenge was to make agent behavior auditable, resumable and governed, rather than a chain of ad hoc prompts.

What I built

  • Designed a persisted workflow state machine (requirements, QA planning, plan approval, development, review, rework, PR publishing, CI monitoring and fixing, human approval) with durable pending-work queues, leases and recovery, so runs survive restarts and resume exactly where they stopped.
  • Built a multi-provider LLM gateway with role-based routing, per-role rate limiting, provider fallback with circuit breaking, automatic continuation of truncated responses, health monitoring with operational alerts, and full audit with optional payload archival.
  • Integrated Cerberus, an in-process prompt-safety layer driven by shared detection rules that blocks secrets, credentials and personal or payment data before any request leaves the trust boundary.
  • Implemented an MCP server exposing least-privilege tools for work-item reading and tagging, branch, file and pull-request operations, code-graph queries, vector memory and test execution, with role-scoped permissions and audit records enforced at runtime.
  • Created a Roslyn-based indexer that emits an immutable, commit-addressed structural graph of projects, types, members and references, giving the planner and dev agent precise code context instead of raw file dumps.
  • Isolated code execution in a container-based workspace worker that clones, edits, builds and tests repositories, plus a browser test runner producing evidence artifacts, all behind approval gates for risky commands.
  • Shipped a real-time portal showing the lifecycle, live agent activity, gates and run history, with approvals and clarifications available both in the portal and as chat cards.

Key decisions and why

01

The agent is not the model, and the gateway is not the agent

Workflow state, tools and memory live in the orchestrator and agent runtimes; the gateway only provides safe, audited, routed model access and knows nothing about work items or branches. This keeps each concern replaceable and the gateway reusable by any caller.

02

Durable, leased state machine instead of in-memory loops

LLM runs are long, flaky and expensive. Persisting the phase, pending work and execution lease allows recovery after crashes, scheduled retries on provider failures, and parking runs that need a human decision, then resuming the same transcript and plan.

03

Prompt safety as an in-process pre-flight gate

Scanning with Cerberus before routing means sensitive content never reaches an external provider, with no extra network hop. Rules are shared data, so they evolve without redeploying the gateway.

04

Separate roles and providers for builder and reviewer

An adversarial reviewer on a different role, and preferably a different model, reduces correlated blind spots. Review findings flow back to the dev agent as structured rework, with a bounded number of iterations before escalating to a human.

05

Tools enforced by the MCP server, not by prompt text

Permissions per agent role, protected-branch rules and branch naming are validated server-side, and no generic mutation tool exists. A compromised or confused prompt cannot exceed its authority.

06

Immutable commit-addressed code graph

Stable symbol identities (hash of repo, qualified name and signature) and per-commit generations make context retrieval reproducible and cheap, and let multi-solution repositories merge into a single index.

Tech stack

Languages
C#TypeScript
Backend
.NET 10ASP.NET CoreRoslyn
AI
Model Context Protocol (MCP)OpenAIAnthropicSelf-hosted LLMEmbeddings (OpenAI / Voyage)
Data
QdrantCouchbase
Cloud
Azure Blob StorageAzure Key Vault
DevOps
Azure DevOpsDocker
Other
Microsoft Teams bot
Frontend
ReactSignalR
Testing
Playwright

Project story

Open fullscreen

Project story

Agentic Software Delivery Platform

Rotate your phone or open fullscreen to watch the story.

Open fullscreen

Click or use the arrow keys to step through