M03
Expansion and evaluation: multiple workflows with distinct roles, stronger prompts and routing, a challenge or synthesis behavior, and real evaluation evidence.
Purpose
This review is where the project becomes a deliberate learning system rather than a working prototype. Two things change: your architecture becomes genuinely multi-workflow, and you start producing evidence about whether the system is any good.
Evaluation evidence is the part most often underestimated. Begin collecting it before this review, not during the week of it.
What to Add
- Multiple workflows or agents with distinct roles. Each should have a responsibility you can state in one sentence.
- Better prompts and routing logic. Prompts should be mode-aware: the system prompts differently when quizzing than when synthesizing.
- A challenge, quiz, synthesis, or application workflow. See Workflow Patterns You Can Build for candidates.
- Evaluation evidence. Test conversations, scoring, critiques, or error analysis. Show what you measured and what it told you.
- Improvements based on prior feedback. Point explicitly to what changed since M02 and why.
What to Show at the Review
- Your multiple workflows and how they are orchestrated.
- Stronger prompts, with before-and-after where you revised them.
- Challenge or quiz behavior running live.
- Your evaluation evidence and what you concluded from it.
- Your improvement plan for the final review.
Expected Outcome
A more capable and more deliberate learning system.
How This Review Is Evaluated
Added workflows, stronger learning behaviors, evaluation approach, response to feedback, and technical improvement. This review contributes 25% of your project grade.
Against the capability checklist, you should be at: architecture stable; chat workflow, materials ingestion, and retrieval/memory all improved; at least 3 to 4 workflows; several learning behaviors; an evaluation workflow that is working and informative; a stronger demo; and clear improvement over M02.
A Reminder on Scope
Resist adding workflows for their own sake. Three or four workflows that reliably produce strong learning behavior score better than seven that half-work. If something is fragile, fix it rather than building alongside it.
Submission Guidelines
(The details for submitting this milestone have not been posted, yet. Please check back.)