โ† Back to Projects

FlowForge

AI Engineering

Autonomous Task Decomposition & Execution Engine

A six-stage agent pipeline that takes a single project brief and turns it into a finished deliverable: analyzing the task, planning the implementation, executing it, testing the result, independently reviewing it, and finalizing the output.

CrewAIAnthropic APIPythonMulti-agent
FlowForge multi-agent pipeline interface

01

Problem

What problem were you solving?

A single generalist LLM call struggles with multi-stage work: it tends to blend planning, research and writing into one pass, with no real checkpoint for quality along the way.

02

Solution

What did you build?

A pipeline of specialized roles, each with a narrow job and a clear handoff to the next, so the system can catch its own mistakes before they reach the final output.

03

Key Features

๐Ÿ”

Analyst

Breaks down the incoming task and scopes what needs to happen.

๐Ÿ“

Architect

Plans the implementation approach before any work begins.

โš™๏ธ

Implementer

Executes the plan and produces the actual work.

๐Ÿงช

Tester

Runs the result against the task's requirements.

โœ…

Reviewer

Independently reviews the finished artifact for quality.

๐Ÿ“ฆ

Finalizer

Assembles and packages the reviewed work into the deliverable.

04

Architecture

How does it work?

Project BriefNatural-language input
โ†’
AnalystScope & requirements
โ†’
ArchitectImplementation plan
โ†’
ImplementerExecutes the plan
โ†’
TesterValidates the result
โ†’
ReviewerIndependent review
โ†’
FinalizerPackaged deliverable

Each stage hands off to the next with its own scope and tool permissions. If the Tester or Reviewer flags an issue, only the affected stages rerun rather than the whole pipeline. Once review passes, the Finalizer assembles and packages the deliverable.

05

Tech Stack & Tools

ComponentPurpose
CrewAIRole-based multi-agent orchestration
Anthropic APIReasoning across all six agent roles
PythonPipeline and agent implementation
Structured outputQuality-threshold exit conditions

06

Technical Decisions

Why these technologies?

Why CrewAI?

Simplifies defining distinct agent roles and letting them hand off work in order, instead of coordinating everything from one script.

Why Anthropic?

Long context and reliable instruction-following made it well suited to research, drafting and review inside the same pipeline.

Why multiple agents?

Separate stages provide independent verification, isolated tool permissions, and stage-level recovery instead of relying on one continuous agent context.

Why quality thresholds?

Gives the pipeline an explicit exit condition, so it iterates until output clears a bar instead of stopping after a fixed number of passes.

07

Results

Benchmarked head-to-head against a baseline system

100%task success: 3/3 benchmark tasks
9.5/10output quality, equal to baseline
100%defect detection: 2/2 injected defects
33%of pipeline rerun on recovery
5.27×latency overhead on simple tasks
5.16×cost overhead on simple tasks

FlowForge matched the baseline on task success, output quality and defect detection. Its recovery loop is also efficient: a failure re-executes only 2 of 6 pipeline stages, not the whole run. The tradeoff is real: roughly 5× the latency and cost of a single-agent baseline on simple tasks, the price of routing work through multiple specialized agents instead of one.

08

Challenges

What difficulties did you face?

  • Designing handoffs between stages without losing the original task's intent
  • Tuning the review threshold to catch real issues without looping forever
  • Keeping each stage grounded in the plan instead of letting it improvise

09

Limitations

What could be better

  • No cap on review iterations for a brief that never quite clears the bar
  • No memory carried between separate brief runs
  • Context quality depends on whatever sources are reachable at run time

10

What I Learned

  • Specialization beats a single mega-prompt for multi-stage work
  • The feedback loop is the product, not any one agent
  • Structured handoffs prevent context loss between stages

11

Future Improvements

  • Cap and surface iteration count instead of looping silently
  • Add shared memory across related briefs
  • Support more output formats (slides, structured JSON)