How to Run Shipley Color Team Reviews: A 7-Step Guide for GovCon Teams (2026)
What is a color team review?
A color team review is a structured checkpoint in the proposal lifecycle where a defined group of reviewers assesses the proposal against defined criteria for that stage. Each color marks a different level of proposal maturity, from strategy validation before writing starts through to final executive sign-off before submission.
The framework was popularised by Shipley Associates and is now standard practice across government contracting, though almost every organisation tailors the stage names and the sequence to fit how it actually works.
The seven color team reviews
| Review | When it runs | The question it answers | Who attends |
|---|---|---|---|
| Blue Team | Capture phase, before or at RFP release | Is our capture plan and solution strategy sound, and is every section assigned to a writer? | Capture manager, solution architect, pricing lead, proposal manager |
| Black Hat | Capture phase | What will our competitors bid, and how do we beat it? | Capture, competitive intelligence, SMEs playing the competitor |
| Pink Team | First full draft, commonly 65 to 70% complete | Does every section have real content, and does the story hold together? | SMEs, technical leads, volume leads, proposal manager |
| Red Team | Near-final draft, commonly 90 to 95% complete | How will a government evaluator score this? | Independent reviewers who did not write, compliance specialists |
| Green Team | Alongside or after Red | Is the price right, approved, and consistent with the technical narrative? | Pricing, finance, contracts, executive sponsor |
| Gold Team | After Red and Green findings are resolved, at or near 100% | Is this ready to submit, and did the fixes actually land? | Executives and final approvers |
| White Hat / White Glove | Pre-submission and post-submission | White Glove: are there production defects? White Hat: what do we do differently next time? | Production staff, then the full team for lessons learned |
Few teams run all seven on every bid. Smaller pursuits often collapse Blue into Pink, or skip Green on firm fixed price work with a thin cost narrative. Red Team is the one to protect.
A note on percentages. Completion percentages are useful shorthand and a poor definition. A more reliable test for Pink Team is whether every section of the proposal is addressed with content, an approach, or at minimum a stated intent. If a section is still a heading and a promise, the review will be about the gap rather than the argument.
How to run a Shipley-aligned color team review in 7 steps
1. Define your stages and what each one is allowed to be about
Map each review to specific exit criteria before the schedule is set. Blue Team should end with a validated strategy and an assigned outline. Pink Team should end with confirmation that the story holds together and every requirement has a home. Red Team should end with an evaluator-style assessment documenting strengths, weaknesses and risks.
The most common failure is treating every review the same way. Apply one feedback lens to all of them and your Red Team spends its time catching problems that belonged to Pink Team, while the evaluator-perspective work that only Red Team can do never happens. Fixing strategy at Blue Team costs hours. Fixing it at Red Team costs weeks.
Write the exit criteria down and put them at the top of the feedback form. A reviewer who knows what this session is not about gives better feedback on what it is about.
2. Assign reviewers by stage, and keep Red Team independent
- Blue Team needs capture managers, solution architects and pricing strategists who can validate strategy before writers commit effort.
- Pink Team needs SMEs and technical leads who can judge whether the solution is strong and whether the response addresses what Section L actually asks for.
- Red Team needs reviewers who did not write any of it. This is non-negotiable. If your Red Team is your writing team, you are marking your own homework, and the one review designed to simulate an outside evaluator instead confirms what the insiders already believe.
- Green Team needs pricing, finance and contracts, with enough technical presence to catch a price that no longer matches the solution described.
- Gold Team needs the executives who will sign, reviewing a document that should hold no surprises.
Rotate reviewers across pursuits. The same three SMEs on every Red Team burn out, and burned-out reviewers give shallow feedback fast.
3. Build the compliance matrix before Blue Team, not after
A review only works if everyone is arguing about the same requirement set. Before kickoff, shred the solicitation to extract every shall, must and will statement, and tag each one by section, owner and complexity. That matrix is the reference every reviewer works from for the rest of the cycle.
How the matrix is built determines whether it can carry that weight. Rules-based extraction reads the document by pattern and produces the same output on every run, traceable to the source sentence. Probabilistic extraction reads with judgement and can produce a slightly different set the second time. When a Pink Team reviewer asks whether requirement 3.4.2 is covered, you want to resolve that in seconds against a shared record, not open a debate about whose export is current.
This is also where scope creep in reviews gets prevented. When coverage is visible, reviewers stop spending the session hunting for gaps and start assessing the quality of what is there.
4. Score against Section M, not against taste
Government evaluators do not reward good prose. They record strengths, weaknesses, deficiencies and risks against stated evaluation criteria. Your scoring should mirror that structure, using your solicitation’s own Section M language and its own weighting.
Define your rating scale tightly enough that it means the same thing to every reviewer:
- Strength. Exceeds the requirement and creates identifiable benefit to the government.
- Acceptable. Meets the requirement without standing out.
- Weakness. Falls short of the requirement or introduces uncertainty.
- Deficiency. Fails to meet a material requirement.
- Risk. Could affect cost, schedule or performance if the approach holds.
Then require specificity. “Strengthen this section” gives a writer nothing. “Add quantified past performance showing delivery within 5% of schedule on comparable IDIQ work” gives them a task they can finish tonight.
5. Run the session with a stage-specific form and a hard time box
Unstructured reviews produce unstructured feedback. Give each stage its own form, built from that stage’s exit criteria, so reviewers are not guessing whether to comment on compliance, positioning or commas.
Time-box the session and the response window. Review length should scale with page count and volume structure rather than a fixed rule, but the failure mode is consistent: reviews that stretch produce feedback that arrives too late to incorporate properly. Fix the date the findings are due, work backwards, and hold it.
One more discipline. Separate the finding from the fix. Reviewers document what they observed and why it matters. Writers and the proposal manager decide the remedy. Reviewers who redraft sentences in the margin are doing the writer’s job and skipping their own.
6. Track findings, ownership and resolution status in one place
Feedback that is not tracked is feedback that gets lost. After each session, consolidate everything into a single record: what was raised, by whom, severity, who owns it, and when it is due.
Triage by severity rather than by volume. Compliance gaps and evaluation risks must close before the next gate. Cosmetic suggestions close if time allows. Without that split, teams reliably spend the Friday before submission on formatting while a Section M weakness sits open.
Then verify the fixes rather than trusting the status column. Comparing the pre-review and post-review versions of a document shows you exactly what changed, which is the only reliable way to confirm that a Red Team finding was actually addressed and not just marked closed.
7. Make Gold Team a verification, not a discovery
Gold Team is the last gate before the government reads it. It should be short, and it should be boring. Everything on this list is checkable rather than debatable:
- Every requirement in the compliance matrix has a traceable response.
- Section L instructions are followed, including format, page limits and volume structure.
- Red Team and Green Team findings are closed, with the change verified in the document.
- Required certifications, representations and forms are present and signed.
- Acronyms are defined at first use and used consistently across volumes.
- Readability holds up under time pressure. An executive summary at grade 10 to 12 reads easily. At grade 15 it slows a tired evaluator down, and a tired evaluator is not on your side.
- Watchwords are clear. Unsupported claims like world class, state of the art or best athlete solution, and legally loaded words like expert or expertise, which put you on the hook to prove them after award.
If Gold Team is discovering problems, the earlier gates did not do their job.
Two mistakes that quietly wreck good reviews
Reviewers arriving cold. A reviewer who spends the first hour working out what the solicitation asks for has an hour left for judgement. Send the compliance matrix, the Section M criteria, the win themes and the stage’s exit criteria in advance, and the session starts at the useful part.
Automating the judgement instead of the preparation. Tooling should compress the mechanical work: extracting requirements, tracking coverage, catching passive voice, flagging watchwords, verifying that a change landed. Deciding whether a technical approach beats the competition is not mechanical work, and a scored output that looks authoritative can make a review feel complete when nobody actually made a call.
“It’s based on our philosophy that we don’t think AI can make these judgments for you.”
Fergal McGovern, CPO, VisibleThread
How VisibleThread supports color team reviews
VisibleThread handles the preparation and verification work around each gate so reviewers spend their time on judgement.
Your stages, your names. The Proposal Kanban Board runs on the stages you actually use: initiated, in progress, pink team, red team, gold team, submitted, or whatever your process really looks like. Stages are configured per workspace, so a two-week commercial bid can run three gates while a multi-year federal pursuit runs eight, in the same platform. Teams that do not use a color team approach define their own.
A compliance matrix every reviewer can trust. Shredding runs on pattern matching, not prediction, so the requirement set is identical on every run and every line traces back to its source sentence. Reviewers work from one record rather than several exports.
Section ownership that survives the real world. Assign volumes and sections to named SMEs, mark them in review or complete, and see who is behind while there is still time. Co-editing is supported, and it is not the only option, because co-editing breaks down the moment the reviewer you need is unavailable until Tuesday.
Verification between gates. Doc Compare shows exactly what changed between two versions, with pattern matching that is 100% accurate every time. Run it on your own drafts to confirm Red Team findings actually landed, and on the inbound solicitation to catch every amendment.
“Before VisibleThread, tracking changes between solicitation versions meant manually going back and forth between documents, a slow, high-risk process with no independent safeguard.”
Marieke Bland, Director of Proposals, KMS Solutions
Clarity work done before Red Team, not at it. Interactive Scoring Mode scores grade level, sentence length, passive voice and complexity per section and per author, with ignore terms so unavoidable technical vocabulary does not distort the grade. Watchword lists flag unsupported and legally risky language on rules you define. You arrive at Red Team with the clarity work already done instead of discovering it there.
Reviewers stay in Word. The Word add-in puts the analysis where your writers and reviewers already work, so a review cycle does not depend on everyone adopting a new editor first.
See it on your own solicitation. Book a demo, bring a live RFP, and we will shred it and set up your review stages in the session.
FAQs
What is the difference between a Pink Team and a Red Team review?
Pink Team reviews the first full draft, commonly at 65 to 70% complete, focusing on whether every section has real content and whether the narrative holds together. Red Team reviews a near-final draft, commonly at 90 to 95%, and simulates how a government evaluator will score it. Red Team reviewers should not have written any of the proposal.
What is a Blue Team review?
Blue Team is a capture-phase review, held before or around RFP release. It validates the capture plan and solution strategy, and in many organisations also confirms the proposal outline is complete and every section has an assigned writer. It identifies information gaps, SMEs, discriminators and win themes before drafting begins.
What is a Green Team review?
Green Team reviews and approves pricing. It confirms the price is competitive, internally approved, and consistent with the technical and management narrative. It typically runs alongside or just after Red Team, so its findings can be resolved before Gold Team locks the document.
How long should each color team review take?
It depends on page count and volume structure rather than a fixed rule, so set the deadline for findings first and work backwards. What matters more than duration is that the response window leaves writers real time to act. Reviews that produce feedback too late to incorporate are the most common scheduling failure in the process.
Who should participate in a Red Team review?
Reviewers who did not write any part of the proposal. Include compliance specialists who can verify requirement coverage, and evaluator-minded reviewers who can score against Section M. Independence is the point of the stage. If your writers are your Red Team, the review confirms what the team already believes.
Can color team reviews be combined when timelines are tight?
Blue and Pink are sometimes merged on smaller pursuits, and Green is sometimes skipped on firm fixed price work with a thin cost narrative. Red Team should not be merged, because merging it removes the independent evaluation that gives it its value. When you consolidate, keep separate exit criteria for each assessment area.
How do you score a proposal during a color team review?
Build the rubric from your solicitation’s Section M, matching its evaluation factors and weighting. Define strength, acceptable, weakness, deficiency and risk tightly enough that the terms mean the same thing to every reviewer. Require findings specific enough to act on, with the evidence or requirement each one refers to.
How do you keep reviewers from burning out?
Rotate reviewers across pursuits rather than using the same SMEs every cycle, build recovery time in after major deadlines, and send materials in advance so sessions start on judgement rather than orientation. Reducing the mechanical work, requirement tracking and compliance verification, is what preserves reviewer attention for the parts that need it.