Corporate project · Presented company-agnostic: architecture and decisions only.
Agentic Software Delivery Platform
Persisted state machine coordinating AI agents from backlog item to approved pull request
Status
In production
Role
Lead engineer and architect: workflow state machine, LLM gateway, prompt-safety integration, MCP tool surface, vector memory, indexer design, portal and deployment shape.
Architecture
4 role-specific AI agents
Persisted state machine with leases
Human approval gates
Work item to approved pull request
A ready signal creates a persisted run; the orchestrator dispatches definition, QA planning, plan approval, development, review and rework, then publishes a pull request, watches CI, and waits for human approval.
Click a step to jump to it. Click a component for details.
What it does
An internal platform that turns work-tracking items into reviewed pull requests by coordinating specialized AI agents for requirements, QA planning, development, adversarial review and CI repair. A multi-provider LLM gateway with a prompt-safety layer, an MCP server exposing scoped tools, vector memory and a code-to-graph indexer provide the shared infrastructure, with human approval gates at every critical step.
The problem
Teams wanted LLM-assisted delivery but could not allow agents to talk directly to providers, leak sensitive data in prompts, hold unrestricted repository or backlog access, or lose work when a long-running run crashed. The challenge was to make agent behavior auditable, resumable and governed, rather than a chain of ad hoc prompts.
What I built
Designed a persisted workflow state machine (requirements, QA planning, plan approval, development, review, rework, PR publishing, CI monitoring and fixing, human approval) with durable pending-work queues, leases and recovery, so runs survive restarts and resume exactly where they stopped.
Built a multi-provider LLM gateway with role-based routing, per-role rate limiting, provider fallback with circuit breaking, automatic continuation of truncated responses, health monitoring with operational alerts, and full audit with optional payload archival.
Integrated Cerberus, an in-process prompt-safety layer driven by shared detection rules that blocks secrets, credentials and personal or payment data before any request leaves the trust boundary.
Implemented an MCP server exposing least-privilege tools for work-item reading and tagging, branch, file and pull-request operations, code-graph queries, vector memory and test execution, with role-scoped permissions and audit records enforced at runtime.
Created a Roslyn-based indexer that emits an immutable, commit-addressed structural graph of projects, types, members and references, giving the planner and dev agent precise code context instead of raw file dumps.
Isolated code execution in a container-based workspace worker that clones, edits, builds and tests repositories, plus a browser test runner producing evidence artifacts, all behind approval gates for risky commands.
Shipped a real-time portal showing the lifecycle, live agent activity, gates and run history, with approvals and clarifications available both in the portal and as chat cards.
Key decisions and why
01
The agent is not the model, and the gateway is not the agent
Workflow state, tools and memory live in the orchestrator and agent runtimes; the gateway only provides safe, audited, routed model access and knows nothing about work items or branches. This keeps each concern replaceable and the gateway reusable by any caller.
02
Durable, leased state machine instead of in-memory loops
LLM runs are long, flaky and expensive. Persisting the phase, pending work and execution lease allows recovery after crashes, scheduled retries on provider failures, and parking runs that need a human decision, then resuming the same transcript and plan.
03
Prompt safety as an in-process pre-flight gate
Scanning with Cerberus before routing means sensitive content never reaches an external provider, with no extra network hop. Rules are shared data, so they evolve without redeploying the gateway.
04
Separate roles and providers for builder and reviewer
An adversarial reviewer on a different role, and preferably a different model, reduces correlated blind spots. Review findings flow back to the dev agent as structured rework, with a bounded number of iterations before escalating to a human.
05
Tools enforced by the MCP server, not by prompt text
Permissions per agent role, protected-branch rules and branch naming are validated server-side, and no generic mutation tool exists. A compromised or confused prompt cannot exceed its authority.
06
Immutable commit-addressed code graph
Stable symbol identities (hash of repo, qualified name and signature) and per-commit generations make context retrieval reproducible and cheap, and let multi-solution repositories merge into a single index.
Tech stack
Languages
C#TypeScript
Backend
.NET 10ASP.NET CoreRoslyn
AI
Model Context Protocol (MCP)OpenAIAnthropicSelf-hosted LLMEmbeddings (OpenAI / Voyage)