The New Scrutiny
What changed in the room
Ten years ago, a well-built deck had a structural advantage: management held the information. Board members, investors, and executive stakeholders could probe, but real verification took weeks they did not have. A confident narrative, delivered well, often carried the day on delivery alone.
That advantage is gone. Governance research now documents the shift directly:
- PwC's survey of corporate directors found that 35 percent say their boards have already integrated AI, including generative AI, into their oversight activities, and PwC expects that share to rise. In a companion survey of C-suite executives, 99 percent said boards should be using AI in their oversight role. The gap between what executives expect and what boards currently do is closing from the board side.
- Directors describe using AI to independently benchmark company disclosures against peers, test management's numbers against outside market data, and analyze the historical board materials management itself provided, building longitudinal views management never assembled for them.
- Board portal and governance platforms now offer features that generate director-specific probing questions from the pre-read materials automatically.
Translate that into what happens to your recommendation:
- Your benchmark gets checked in real time. The peer comparison you selected can be re-run, with peers you did not select, before you finish presenting.
- Your numbers get tested against your own history. "This projects 25 percent growth; the last three initiatives you presented projected similarly and delivered 8. What is different?" is now a question anyone can construct in minutes from your old decks.
- Your pre-read gets red-teamed. The two-day pre-read window is enough time for a stakeholder to have AI generate the twenty hardest questions your document raises, complete with the page references.
- Alternatives you did not present get surfaced. "What options did you consider and reject?" now comes with the stakeholder having already generated the option list themselves.
The asymmetry, and how to flip it
Here is the uncomfortable part: challenging a recommendation with AI takes minutes, while building one still takes weeks. The attacker's cost dropped faster than the builder's. That asymmetry is why polished-but-fragile narratives are failing at higher rates in rooms that used to accept them.
The flip is equally available: every tool your stakeholders can point at your recommendation, you can point at it first. The same AI that generates their twenty hardest questions can generate them for you, two weeks earlier, while you still have time to fix what the questions expose. That is the whole strategy of this course. Scrutiny is coming either way. The only variable you control is whether it happens before the meeting, run by you, or during the meeting, run by them.
Why smart people build fragile narratives
Before the method, it helps to know why the flaws in Module 2 are so consistent across industries and levels. They are the signatures of well-documented reasoning patterns, and knowing the pattern makes the flaw easier to see in your own work.
Motivated reasoning. Once a person believes a conclusion, they evaluate evidence for it more leniently than evidence against it. This is not dishonesty. It is the default setting of a mind that has already decided. Every weakness in Module 2 is a shortcut that motivated reasoning makes attractive.
The planning fallacy. Decades of research on forecasting show that people systematically underestimate the time, cost, and risk of their own projects while estimating other people's projects more accurately. The inside view (this specific plan, its specific strengths) crowds out the outside view (how projects of this type usually go). Projections built from the inside view bend upward.
Reference class neglect. The correction for the planning fallacy is to ask how a reference class of similar initiatives actually performed, and to anchor the forecast there before adjusting for specifics. Most executive narratives skip this step entirely. Your stakeholders, with AI, no longer do.
Confirmation in sourcing. When a person chooses which evidence to gather, they gather the evidence most likely to support what they already believe. The vendor's study gets read. The independent analysis with the inconvenient number gets a footnote.
Sunk cost and escalation. Once an initiative is underway, every dollar spent becomes an argument for the next dollar. A plan with no predefined stopping condition converts this bias into policy.
The method in this course is a set of external checks on these internal patterns. It works because it does not rely on you noticing your own motivated reasoning, which is the one thing motivated reasoning prevents.
What survives scrutiny
Recommendations that hold up under AI-enabled pressure share a structural property: their reasoning is visible and already stress-marked. They state their assumptions rather than burying them. They anchor projections to a reference class before adjusting. They show the alternatives considered and the honest reason each lost. They name their risks with responses attached, and they define in advance what evidence would change the recommender's mind.
Notice what that list does not include: certainty. Defensibility does not come from having no weaknesses. It comes from knowing your weaknesses better than your challengers do, and having decided what to do about each one. A leader who says "this assumption is the load-bearing one, here is why I believe it, and here is what we will see within 90 days if I am wrong" is harder to rattle than one defending a claim of confidence.
Check your understanding
- What specific capabilities can stakeholders now apply to your recommendation between pre-read and meeting?
- Why does the attacker-builder cost asymmetry favor whoever stress-tests first?
- Which reasoning pattern produces a projection that bends away from history, and what is its standard correction?
- What is the difference between a recommendation with no visible weaknesses and a defensible one?
Exercise 1: Name your exposure
10 minutesTake your real case. Without deep analysis yet, write quick answers:
- Which single number in your case would you least like a stakeholder to independently verify?
- Which comparison or benchmark did you choose, and what would a differently chosen one show?
- What alternative did you not present, and what is your honest reason?
- What is the reference class for your initiative (the set of similar efforts, inside or outside your organization), and do you know how that class typically performed?
- What would have to be true for you to withdraw this recommendation?
Do not fix anything yet. This is a baseline of where you already suspect the soft spots are. Module 3 tests whether your suspicions are right and finds the ones you missed.
Quick check
Two questions on this module.
The Seven Weaknesses of Executive Narratives
Why the same flaws keep appearing
Strategic recommendations fail under scrutiny in patterned ways. The patterns repeat because they are the shortcuts a smart, honest person takes when they already believe their conclusion. Knowing the seven patterns gives you a checklist for finding them in your own work, where they are hardest to see.
For each pattern below: what it looks like, its common variants, the detection question that exposes it, and the repair.
1. The buried assumption
What it is. The recommendation depends on something never stated, so it is never examined. "The rollout completes in Q2" quietly assumes the two key hires close in Q1, vendor timelines hold, and no community-level disruption intervenes. The narrative presents the conclusion of a chain whose links appear nowhere.
Variants. The dependency chain (A requires B requires C, and only A is stated). The stable-environment assumption (competitors, regulators, and prices hold still). The capacity assumption (the same team that runs today's operation also delivers the initiative, with no degradation to either).
Detection question. "For this claim to be true, what else has to be true?" Ask it three times in a row, chaining downward.
Repair. State the assumption, rate it (Module 3), and if it is load-bearing, give it its own line in the document with the evidence behind it.
2. The chosen baseline
What it is. Every comparison requires choosing what to compare against, and the choice is doing silent work. Growth "versus last year" when last year dipped. Benchmarks against the top quartile when the honest peer set is the median. Costs versus the original budget rather than the revised one. The number is accurate. The frame was selected.
Variants. The favorable period (the start date chosen after the bad quarter). The favorable peer set (competitors selected because they flatter). The favorable metric (revenue growth shown, margin not). The moving target (compared against a plan that was itself revised).
Detection question. "Compared to what, and who chose it?" Then: "What does the comparison show against the most obvious alternative baseline?"
Repair. Show the comparison against the baseline a skeptic would choose, alongside yours, and explain why yours is the right one. If you cannot explain it, use theirs.
3. The hockey stick
What it is. History grows at 6 percent; the projection grows at 25, with the bend arriving conveniently after the decision. The tell is a projection whose driver is the initiative's assumed success rather than any mechanism you can name.
Variants. The delayed bend (flat for two years, then vertical). The compounding assumption (each year's growth rate assumed to hold or rise). The synergy line (an unexplained uplift labeled "synergies" or "network effects"). The best-case-as-base-case (the optimistic scenario presented without the others).
Detection question. "What specifically produces the bend, and when will it first be visible in a number we can observe?" Then: "How did the reference class of similar initiatives actually perform?"
Repair. Anchor to the reference class. Show the base rate for initiatives of this type, then show your adjustment from it and the specific reasons for the adjustment. Present the projection as a range with the driver of each end named.
4. The friendly pilot
What it is. Evidence from a test run under the conditions most likely to make it succeed: the flagship location, the strongest team, the most engaged community, volunteer users. Real result, wrong denominator. The pilot proves the idea can work, then the narrative treats it as proof the idea will work everywhere.
Variants. The volunteer effect (participants who opted in differ from those who will be required). The attention effect (the pilot got leadership focus the rollout will not). The small-n result (a percentage improvement from a base of twelve). The short window (measured before novelty wore off).
Detection question. "How does the pilot population differ from the population I am projecting onto, and in which direction would each difference push the result?"
Repair. Name the differences and discount the projection for each. Where possible, run or cite a second test under less favorable conditions. If you cannot, say the pilot proves feasibility, not scale, and build the ask accordingly.
5. The missing alternative
What it is. The case argues for the recommendation without arguing against the alternatives, including the strongest one: doing nothing, or doing a smaller reversible version first. When no alternative appears, stakeholders correctly suspect the comparison was not run, or was run and lost.
Variants. The straw alternative (options shown only to be knocked down). The false binary (do this or do nothing, with the staged version omitted). The missing opportunity cost (what else the money, people, and attention could do).
Detection question. "What is the strongest honest case for not doing this, and for doing a smaller version, and where in my document does each one lose on the numbers?"
Repair. Build the steelman alternatives (Module 3, step 4) and show why each lost, in numbers where possible. If the staged version is nearly as good, recommend it.
6. The convenient source
What it is. Key evidence comes from a party with a stake in the conclusion: the vendor's ROI study, the consultant hired by the initiative's sponsor, the market analysis commissioned to support a decision already leaning. Also in this family: the three-year-old market study still cited because updating it might change the answer.
Variants. The interested party (vendor, sponsor, consultant). The stale study (accurate when produced, unexamined since). The single source (one report doing all the evidentiary work). The survivor sample (a benchmark built only from the successes).
Detection question. "Who produced this, what did they want to be true, when is it from, and what would an independent source with current data show?"
Repair. Triangulate: find a second source with no stake. Where you cannot, say so, and treat the claim as an assumption rather than evidence.
7. The all-or-nothing ask
What it is. Full commitment requested up front, no staged decision points, no defined conditions under which the initiative stops. The absence of kill criteria signals that the recommender has not seriously imagined being wrong, and it converts every future sunk cost into an argument for continuing.
Variants. The urgency wrapper (the full commitment justified by a time pressure that carries no evidence). The lock-in structure (contracts or leases that make stopping expensive by design). The missing tripwire (no stated evidence that would trigger a pause).
Detection question. "What would we observe, and by when, that would cause us to stop, and what does stopping cost at that point?"
Repair. Stage the ask with decision gates. Define tripwires and responses (Module 3, step 5). Present them as part of the recommendation, not as a concession.
Quick-scan table
| Pattern | Tell | Detection question |
|---|---|---|
| Buried assumption | A conclusion with no visible chain | What else has to be true? |
| Chosen baseline | A flattering comparison | Compared to what, chosen by whom? |
| Hockey stick | Projection bends away from history | What produces the bend, and when is it visible? |
| Friendly pilot | Best-case evidence generalized | How does the pilot differ from the rollout population? |
| Missing alternative | No do-nothing or staged case | Where does the strongest alternative lose? |
| Convenient source | Interested or stale evidence | Who produced it, wanting what, when? |
| All-or-nothing ask | No gates, no tripwires | What would make us stop, and when? |
Safe case 1: Bluewater Coffee Roasters
Read this fictional recommendation. It contains at least five of the seven weaknesses, planted deliberately. Find them before the reveal in Module 3.
Bluewater Coffee Roasters: Retail Expansion Recommendation (excerpt from CEO memo to the board)
Bluewater operates 38 cafes across the Pacific Northwest with $46M in annual revenue. I recommend the board approve an $18M investment to open 12 new locations across two new metro markets over 18 months, alongside launch of the Bluewater mobile app.
The specialty coffee retail segment continues to grow, with leading chains posting 22 percent annual revenue growth. Our own app pilot at the Portland flagship generated a 31 percent increase in repeat visits, demonstrating strong customer appetite for digital engagement across our footprint. A market study by Cascade Growth Partners, who supported our original expansion in 2023, identifies both target metros as high-opportunity markets.
We project the new locations reach $14M combined annual revenue by year three, reflecting 25 percent year-over-year growth after opening. Company revenue has grown steadily at 5 to 7 percent annually over the past four years, and this initiative moves us decisively beyond that plateau. Speed matters: delaying entry risks ceding both markets to competitors. I recommend we commit to all 12 locations now to secure favorable lease terms, with construction beginning in Q1.
Safe case 2: Harborview Community Services
A second fictional case from a different sector, so you can see that the patterns are not about retail. This one is not worked in Module 3; it is yours to practice on in Exercise 2 and again in Module 3.
Harborview Community Services: Case Management Platform Recommendation (excerpt from COO memo to the executive committee)
Harborview serves about 9,000 clients a year across 14 sites with a $31M operating budget. I recommend we approve a $2.4M, three-year contract with Meridian Systems to replace our case management platform across all 14 sites, with cutover in a single weekend next spring.
Meridian's implementation study projects a 30 percent reduction in caseworker documentation time, based on results at comparable agencies. Our three-month pilot at the Eastside site, led by our most experienced program manager, showed a 26 percent reduction and strong staff satisfaction. Caseworker turnover, currently 34 percent annually, is projected to fall to 20 percent within two years as documentation burden drops, saving an estimated $900K a year in recruiting and training costs.
The current platform's vendor has announced end of support in 20 months, so a decision is needed now. A phased rollout was considered and rejected because running two systems in parallel would confuse staff. I recommend the full three-year commitment to secure Meridian's multi-year pricing.
Exercise 2: Three audits
20 minutesAudit 1, Bluewater (6 minutes): List every weakness you can find in the Bluewater memo. Name which of the seven patterns each one is. Aim for five or more.
Audit 2, Harborview (7 minutes): Same exercise. This case has at least six of the seven. Note that one of them is well disguised as prudence.
Audit 3, your real case (7 minutes): Run the seven patterns as a checklist against your own recommendation. For each, write "clear," "present," or "not sure." Be honest on "not sure": those are usually "present." For each "present," write the detection question's answer in one line.
Keep all three lists. Module 3 gives you the Bluewater reveal, the Harborview answer key at the end, and the repair method for both.
Quick check
Two questions on this module.
The Stress Test
The method: strip, rate, attack, compare, break
The stress test is five steps run in order, each with AI as your red team, followed by a rebuild. It takes 60 to 90 minutes for a significant recommendation, which is why it belongs a week before the meeting, not the night before. This module runs the full method on the Bluewater case; the toolkit appendix carries the prompts for your real work.
Step 1: Strip it to the claim stack
Reduce the narrative to its logical skeleton: the recommendation on top, the three to five key claims holding it up, and under each claim, the assumptions it rests on. Prose hides logic. Stacks expose it.
The Bluewater memo stripped:
Reveal the stripped claim stack
Recommendation: Invest $18M in 12 locations plus the app, all at once, starting Q1.
Claim 1: The market opportunity is large and growing. Rests on: segment growth applies to us (benchmark is "leading chains," pattern 2, the chosen baseline, and a survivor sample, pattern 6); the Cascade study is current and independent (commissioned by the sponsor of the last expansion, pattern 6, the convenient source; its date is unstated).
Claim 2: Customer demand for our model is proven. Rests on: the flagship pilot generalizes to 12 new locations in metros where the brand is unknown (pattern 4, the friendly pilot); "repeat visits" translates into revenue (unstated, pattern 1).
Claim 3: The financial projection is achievable. Rests on: 25 percent growth from a company that has grown 5 to 7 percent for four years, with no stated mechanism for the bend (pattern 3, the hockey stick), plus unstated assumptions about hiring, construction timelines, and existing-store performance not degrading while attention shifts (pattern 1).
Claim 4: Committing fully now is the right structure. Rests on: the urgency claim ("ceding markets") which carries no evidence, the lease-terms claim which is unquantified, and the absence of any staged alternative or stopping condition (patterns 5 and 7).
If your Audit 1 list caught five of those, you found the plants. The stack format is what makes them findable in your own work, where motivated reasoning hides them from the prose reader you become.
A practical note on stripping your own document: do it with AI first (prompt in the appendix), then do it yourself without looking at the AI's version, then compare. The assumptions that appear on only one list are the ones most worth examining.
Step 2: Rate the assumptions
Not every assumption deserves attack. Rate each on two axes: how load-bearing (does the recommendation survive if this is wrong?) and how uncertain (how strong is the actual evidence?). Your attack list is the quadrant that is both load-bearing and uncertain.
| Bluewater assumption | Load-bearing | Uncertainty | Attack? |
|---|---|---|---|
| Pilot generalizes to new metros | Fatal if wrong | Thin (one site, flagship) | Yes, first |
| 25 percent growth rate | Fatal if wrong | None stated | Yes, first |
| Cascade study is current and independent | Damaged if wrong | Thin (interested party, undated) | Yes |
| Lease terms require full commitment now | Survives if wrong (structure changes, not opportunity) | None stated | Later |
| Existing stores hold performance during expansion | Damaged if wrong | Not addressed | Yes |
| Hiring and construction cadence achievable | Damaged if wrong | Not addressed | Yes |
Spend your scrutiny where wrongness is fatal. The lease-terms claim is uncertain but shapes the structure of the ask, not the opportunity itself, so it ranks below the two fatal assumptions.
Step 3: Attack the evidence
For every piece of evidence supporting a load-bearing claim, run five questions:
- Source: who produced this, and what did they want to be true?
- Age: when is this from, and what has changed since?
- Denominator: what population does this actually describe, and is it the population I am projecting onto?
- Counterfactual: compared to what? What would this number look like under the do-nothing case, or with a different baseline?
- Survivors: am I looking at the winners only? (The "leading chains at 22 percent" benchmark excludes every chain that expanded and shrank.)
Worked on Bluewater's two fatal assumptions:
The pilot (31 percent increase in repeat visits).
- Source: internal, produced by the team proposing the app. Wanted it to work.
- Age: unstated. If the pilot was recent, novelty effects are still in the number.
- Denominator: one flagship store in the home market with the highest brand awareness in the chain. The projection is onto 12 stores in two metros with zero brand awareness. Every difference pushes the result down.
- Counterfactual: what did repeat visits do at the other 37 stores over the same period? If they rose 15 percent on their own (seasonality, a promotion), the app's effect is 16 points, not 31.
- Survivors: not applicable to a single site, but the choice of the flagship is itself a selection.
The 25 percent growth rate.
- Source: management projection. No external anchor.
- Age: not applicable.
- Denominator: new locations only, which is fair, but the claim that it "moves us beyond the plateau" mixes new-store growth with company growth.
- Counterfactual: what is the reference class? New-location revenue ramps for specialty coffee chains entering new metros. If that class typically reaches 60 to 70 percent of mature-store revenue by year three, the $14M figure needs to be shown against it.
- Survivors: the 22 percent benchmark comes from "leading chains," which is a survivor sample by definition.
Add a sixth question when the evidence is a projection rather than a fact:
- Reference class: how did initiatives of this type, in this sector, at this scale, typically perform? Anchor there first, then adjust with named reasons.
This is where AI earns its place. Hand it your evidence summary and have it run the six questions adversarially. It will not know your industry's ground truth, but it is relentless at spotting which claims rest on which sources, and it does not share your motivation to go easy. It is also good at suggesting what the reference class might be, which you then verify.
Step 4: Compare against the steelman alternatives
Build the strongest honest version of at least two alternatives: the do-nothing case and the strongest different approach. For Bluewater:
(a) Do not expand. Invest the $18M in same-store growth and the app across all 38 existing cafes. Steelman: the app pilot's own logic says digital engagement lifts repeat visits; applying it across 38 known stores with existing brand awareness is lower risk than 12 unknown ones. If the app lifts existing-store revenue 8 percent, that is $3.7M a year on $46M with no construction risk.
(b) Stage it. Open 3 locations in one metro with defined success gates, then decide on the rest. Steelman: it tests the pilot-generalization assumption with real data at a quarter of the capital. If the first 3 hit projections, the remaining 9 proceed with a demonstrated model and better lease negotiations (a proven concept commands better terms than a promise). If they miss, $13.5M is preserved.
The test your recommendation must pass: state, in numbers where possible, why it beats each steelman. If the staged alternative is nearly as good with a fraction of the downside, the honest recommendation may be the staged one, and discovering that before the board does is the difference between leading the meeting and losing it. Stakeholders with AI will generate these alternatives. The only question is whether your document already contains your answer.
Step 5: Break it, then set the tripwires
Run a pre-mortem: it is 18 months from now and the initiative failed. Write the three most plausible failure stories, specifically, with causes.
The pre-mortem is one of the best-evidenced tools in this course. Research on "prospective hindsight" found that imagining an outcome has already occurred, and then explaining it, substantially increases the number and specificity of causes people can identify, compared with asking what might go wrong. The technique works because it converts the question from "will this fail?" (which motivated reasoning answers "no") to "why did this fail?" (which the mind treats as a puzzle to solve).
Bluewater's three failure stories:
- The pilot did not travel. App adoption in the new metros ran at a third of the Portland rate because nobody knew the brand. Repeat-visit lift was 6 percent, not 31. Year-three revenue reached $8M against a $14M projection.
- The cadence broke. Twelve openings in 18 months meant one every six weeks. The third location opened four months late when the general manager hire fell through; construction on locations five through eight stacked up behind it. Costs ran 30 percent over and the opening schedule slipped a year.
- The core business paid for it. Leadership attention, the best store managers, and marketing budget all moved to the new metros. Same-store sales in the original 38 fell 4 percent, wiping out the year-one contribution from the new stores.
Then convert each into two artifacts:
- A tripwire: the earliest observable evidence that this failure story is beginning.
- A response: what happens when the tripwire trips.
| Failure story | Tripwire | Response |
|---|---|---|
| Pilot did not travel | First 3 locations below 60 percent of projected revenue at month 6, or app adoption below 40 percent of Portland rate at month 3 | Pause further openings; revisit the model with actual data before releasing capital for locations 4 to 12 |
| Cadence broke | Any location more than 6 weeks behind schedule, or GM role unfilled 60 days before planned opening | Slow the cadence to one opening per quarter; do not start construction on a location without a signed GM |
| Core business paid | Same-store sales in original 38 down more than 2 percent year over year for two consecutive quarters | Reassign a named executive to core operations; freeze new-market marketing spend until recovery |
Tripwires plus responses are your kill criteria, and they repair pattern 7. Presenting them does not weaken your ask. It is usually the single strongest credibility move in the room, because it demonstrates you have imagined being wrong and priced it.
Rebuild: the stress-marked narrative
The output of the five steps is not a pile of doubts. It is a stronger document:
- Claims stated with their assumptions visible, the load-bearing one flagged
- Projections anchored to a reference class, with the adjustment and its reasons shown
- Evidence that survived the attack, weak evidence replaced or caveated honestly
- Alternatives shown, with the numeric reason each lost
- Risks named with tripwires and responses attached
- An ask restructured where the stress test demanded it (for Bluewater, staged)
Here is the Bluewater recommendation rebuilt. Compare it to the original.
Bluewater Coffee Roasters: Staged Retail Expansion (rebuilt excerpt)
Recommendation. Approve $6M to open 3 locations in the Boise metro over 9 months, with the app deployed first across all 38 existing cafes. Approval of the remaining $12M for 9 further locations is a separate board decision at month 12, conditional on the gates below.
The case. New-metro specialty coffee entries in our reference set reach 55 to 70 percent of mature-store revenue by year three; we are projecting the midpoint, 62 percent, or $9.3M for 12 locations, not the $14M in the original plan. The app is the load-bearing assumption: its Portland pilot lifted repeat visits 31 percent, but that was our highest-awareness store. Deploying it across all 38 existing cafes first tests whether the effect holds in average stores, and we will have that data by month 4.
What we tested and rejected. Do nothing but the app: projected $3.7M annual uplift at low risk, and it is now part of this plan rather than an alternative to it. Full 12-location commitment: rejected because the pilot-generalization risk is fatal if wrong and untested; staging preserves $12M against that risk at a cost of an estimated 3 to 5 percent on lease terms.
Tripwires. First 3 locations below 60 percent of projection at month 6: pause. Any opening 6 weeks late or a GM unfilled at 60 days out: slow the cadence. Same-store sales down more than 2 percent for two quarters: freeze new-market spend.
What would change my mind. If the app deployment across existing stores shows a repeat-visit lift under 10 percent by month 4, I will recommend the board not release the second tranche.
The rebuilt version asks for a third of the money, projects two-thirds of the revenue, and is far more likely to be approved, because every question the board would have asked is already answered in the document.
Harborview answer key
For your Audit 2. The Harborview memo contains:
- Convenient source: Meridian's own implementation study supplies the 30 percent figure.
- Friendly pilot: Eastside site, led by the most experienced program manager. Every difference from the other 13 sites pushes the result down.
- Hockey stick: turnover from 34 percent to 20 percent in two years, with documentation burden as the only named mechanism, and no reference class for whether platform changes move turnover at all.
- Buried assumption: the $900K savings depends on turnover falling, which depends on documentation time falling, which depends on the pilot generalizing, and on caseworkers leaving mainly because of documentation (unstated and unverified).
- Missing alternative, disguised as prudence: the phased rollout was "considered and rejected" in one sentence with an unquantified reason. The steelman (phase by region, keep parallel systems for 30 days) was not built. A single-weekend cutover across 14 sites is the higher-risk option presented as the safer one.
- All-or-nothing ask: the full three-year commitment, justified by multi-year pricing, with no gates and no tripwires. The end-of-support deadline is real but 20 months away, which is time enough for a staged approach.
Six of seven. The chosen baseline is the one largely absent, though "comparable agencies" in Meridian's study is a baseline chosen by the vendor.
Exercise 3: Stress-test your real case
25 minutes, first passRun steps 1 and 2 on your real recommendation now:
- Strip it to the claim stack: recommendation, claims, assumptions under each. Do it with AI, then by hand, then compare. (12 minutes)
- Rate the assumptions and circle your load-bearing-and-uncertain quadrant. (5 minutes)
- Write your attack plan: which two assumptions get steps 3 through 5 first, what the reference class for your projection is, and what evidence you will need to pull. (5 minutes)
- Draft one tripwire for the failure story you already fear most. (3 minutes)
The full run on a real case takes the 60 to 90 minutes noted above. Schedule it now for a real block this week. The toolkit appendix contains every prompt.
Quick check
Two questions on this module.
Preparing for Sharper Questions
The question bank: generate their prep before they do
Module 1 established that stakeholders can have AI generate their hardest questions from your pre-read. Your move is to generate that question bank first, from personas that match how real scrutiny actually arrives. Four personas cover most rooms.
The verifier. Independently checks your numbers against public benchmarks and your own history. Motivated by accuracy; annoyed by round numbers with no source. Asks: "Your last two initiatives projected above 20 percent and delivered under 10. Walk me through why this projection is different." Also asks: "Where does the 22 percent benchmark come from, and who is in the sample?"
The capital allocator. Treats your ask as competing with every other use of the money. Motivated by return on the whole portfolio; annoyed by asks that assume the money is already theirs. Asks: "Why is this the best $18M we can spend, versus the alternatives you have not shown me?" Also asks: "What is the return if we do half of this?"
The operator. Attacks execution, not strategy. Motivated by having run something this size; annoyed by plans that assume capacity exists. Asks: "Twelve locations in 18 months means an opening every six weeks. Show me the hiring and construction plan that supports that cadence." Also asks: "Who runs the existing 38 stores while your best managers open the new ones?"
The historian. Remembers everything the organization has tried. Motivated by not repeating mistakes; annoyed by plans that ignore the last attempt. Asks: "How is this different from the 2021 expansion, and what specifically did we learn from it that changed this plan?" Also asks: "We said the same thing about urgency last time."
Two supplementary personas for specific rooms:
- The regulator (for anything touching compliance, safety, or licensure): "What is the compliance exposure if this goes wrong, and who signs off before go-live?"
- The affected party (for anything touching staff, clients, or customers): "What does this cost the people who have to live with it, and did you ask them?"
Run your rebuilt document through all four (prompts in the toolkit), plus any supplementary persona that fits your room. Merge the output into a single ranked question bank: the twenty hardest, ordered by how much damage an unprepared answer would do.
The answer architecture
For each question in the bank, prepare the answer in a fixed shape: the direct answer in the first sentence, the evidence in the second, the implication for the decision in the third. The shape matters under pressure because pressure produces rambling, and rambling reads as evasion even when it is not.
Weak: "That is a great question, and there are several factors to consider around the growth rate..."
Strong: "The projection assumes 62 percent of mature-store revenue by year three, the midpoint of our reference set; our original draft said 25 percent growth and we revised it down. The difference from history is the app-driven repeat rate, which is the load-bearing assumption in this case, and it is exactly what the month-4 deployment across existing stores tests. If that lift comes in under 10 percent, I will recommend you not release the second tranche."
The strong version does something the weak one cannot: it demonstrates the stress test happened. That demonstration, repeated across a few hard questions, changes the meeting's dynamic from prosecution to collaboration, because stakeholders stop hunting for the weakness you are hiding once it is clear you are not hiding any.
Three more shapes worth having ready:
- The concession: "You are right that X. Here is what I have done about it, and here is what it changes." Conceding a real point quickly buys more credibility than defending it.
- The reframe, used sparingly: "The question assumes Y; the data shows Z, and that changes the answer." Only when the premise is actually wrong. Used on a correct premise, it reads as evasion.
- The bridge: "That connects to the load-bearing assumption. Here is how." Used to pull a scattered question back to the thing that matters.
The question you cannot answer
There will be one. The protocol has three rules:
- Never improvise a number. A made-up figure in a board meeting is the single most expensive sentence you can say, because the room can now check it before the meeting ends.
- Bound what you do know. "I do not have the churn figure by market. What I can tell you is the blended rate and its trend, and that the market-level split has not moved the blended number more than two points historically."
- Commit to a date, not an intention. "You will have the full breakdown by Thursday" beats "we will look into that." Then deliver Thursday, because the follow-through is itself evidence for everything else you claimed.
Reading the room
Sharper questions arrive in a few recognizable forms, and each wants a different response.
- The genuine probe. The stakeholder wants to understand. Answer in the architecture, then stop. Over-answering a genuine probe wastes their goodwill.
- The test. The stakeholder knows the answer and wants to see if you do. Answer directly and briefly. Do not explain what they already know.
- The signal to the room. The question is really a statement to other stakeholders. Answer the substance, then let it go. Do not argue the statement.
- The pile-on. Several stakeholders following one hard question. Pause, name the underlying concern ("the theme here is whether the pilot generalizes"), answer that once, and offer the detail offline.
The murder board
The final preparation step, two or three days out: a live simulated hostile Q&A. With a colleague if you can get one; with AI in persona if you cannot, and the AI version has one real advantage: it does not soften out of politeness.
Run it in character, out loud, answering in real time. Rules:
- No restarting answers.
- No checking notes for the first response.
- Every answer scored afterward against the architecture (direct answer first? evidence? implication?).
- At least one round where the questioner follows up on the weakest part of every answer, because follow-ups are where rehearsed answers collapse.
Two rounds of twenty minutes produces more improvement than any amount of silent review, because the failure mode in the room is not knowledge, it is retrieval under pressure, and retrieval is trainable.
If you use a colleague, give them the question bank and the personas, and ask them to be harder than the real room. If you use AI, the murder board prompt in the appendix runs the whole session, including the scoring.
Exercise 4: Build your bank and run one round
20 minutesOn your real case:
- Generate the question bank from all four personas, plus any supplementary persona your room needs (prompts in the toolkit). Merge and rank the top ten. (7 minutes)
- Write architecture-shaped answers to the three hardest. (7 minutes)
- Identify the one question you currently cannot answer well, and write the bounded response plus the commit-by date for closing the gap. (3 minutes)
- Run a five-question murder board round with AI, out loud, and score it. (3 minutes)
Schedule the full murder board for two to three days before your actual presentation. Put it on the calendar now.
Quick check
Two questions on this module.
The One-Pager, the Pre-Read, and Life After Approval
The decision one-pager
The most valuable artifact for recurring stakeholder relationships: a single page that travels with your recommendation.
Recommendation: one sentence, including the ask and its structure. The case: the three to five claims, one line each. Load-bearing assumption: named explicitly, with the evidence behind it and the tripwire that tests it. Alternatives considered: each with the one-line reason it lost. Risks and responses: top three, each with its tripwire. What would change my mind: the evidence that would cause you to withdraw or restructure the recommendation.
That final line is the one leaders resist and the one that buys the most trust. It converts your recommendation from a position to be defended into a decision process stakeholders can see, and decision processes are what boards and investors actually evaluate over time. Anyone can be right once. The one-pager is how you become someone whose recommendations get approved faster with each cycle.
Write it last, because it is the distillation of everything the protocol surfaced. If any line of it is hard to write, that difficulty is your final warning.
Designing the pre-read for AI-armed readers
Your pre-read will be red-teamed. Design it knowing that.
- Lead with the one-pager. The first page is the decision summary. Readers who use AI to generate questions will generate them from a document that already contains your answers.
- Put the assumptions where the AI will find them. A section titled "Assumptions and evidence" gets extracted cleanly. Assumptions scattered through prose get reconstructed by the reader's AI, in whatever form it chooses.
- Show the alternatives with numbers. A reader's AI asked "what alternatives were not considered" will find that they were.
- Include the reference class. A projection with its base rate shown is one the verifier persona cannot easily attack.
- State the tripwires as commitments. "We will report against these at each quarterly meeting" tells the reader you expect to be held to them.
- Keep the deck subordinate. Slides are for the meeting. The pre-read is the document of record, and it should stand without the presenter.
A useful check: run the reader's own move against your pre-read. Hand the document to AI with the prompt "You are a skeptical board member. Generate the twenty hardest questions this document raises." If more than a few of the questions are not already answered in the document, you are not done.
Proactive disclosure
The instinct is to leave the weak spot out and hope nobody asks. In an AI-armed room, somebody will. The better move is to disclose it first, framed with your response.
"The weakest part of this case is the pilot's generalizability. It ran in our highest-awareness store. That is why the plan deploys the app across existing stores first and gates the second tranche on that result."
Disclosure takes the weakness off the table as a discovery and puts it on the table as evidence of judgment. Stakeholders who came prepared to expose it find it already exposed, and they redirect their attention to the substance.
Life after approval: reporting against tripwires
Approval is the beginning of the credibility cycle, not the end. The tripwires you set become the structure for every update that follows.
- Report status against each tripwire, every cycle. Green, watch, or tripped, with the number. A tripwire nobody reports on was a decoration.
- When a tripwire trips, say so first. Before the stakeholder finds it. The response you defined in advance is now the recommendation, and it arrives as a plan rather than a confession.
- Log the decision and the outcomes. If you built a second brain, the one-pager goes into decisions.md with its "revisit if" lines set to the tripwires. Six months later, the historian persona is you, with the record to prove it.
This is the quiet compounding at the center of the course: a recommendation that was stress-tested, staged, and reported against its own tripwires makes the next recommendation from the same person easier to approve. Stakeholders learn that your numbers survive checking, and they check less.
Exercise 5: Draft your one-pager
12 minutesOn your real case, using the template above:
- Fill every section. Where you lack the content, write [NEEDED: description] rather than inventing it. (8 minutes)
- Write the proactive disclosure sentence for your weakest spot. (2 minutes)
- Note which tripwires you will report against, and at what cadence. (2 minutes)
Count the [NEEDED] markers. Each one is a task for the full stress test you scheduled in Module 3.
Quick check
Two questions on this module.
The Pre-Presentation Protocol
The repeatable approach, on a timeline
Everything in this course compresses into a protocol you run before any high-stakes recommendation. Anchor it to the meeting date:
T-minus 7 days: the stress test. Strip, rate, attack, compare, break. Rebuild the document with assumptions visible, projections anchored, alternatives shown, and tripwires set. (60 to 90 minutes)
T-minus 4 days: the question bank. Four personas plus any supplementary ones, merged and ranked. Architecture-shaped answers to the top ten. Identify the can't-answer gaps and start closing them. (45 minutes)
T-minus 3 days: the pre-read check. Run the skeptical-reader prompt against your pre-read. Fix any question the document does not already answer. (20 minutes)
T-minus 2 days: the murder board. Live, in character, out loud, scored, with follow-ups. Fix the answers that collapsed. (40 minutes)
T-minus 1 day: the one-pager. Write it last. Add the proactive disclosure. If any line is hard to write, that difficulty is your final warning. (20 minutes)
Day of: nothing new. No new numbers, no new slides, no new arguments after the murder board. Late additions are untested by definition, and untested claims are where prepared presenters get hurt.
T-plus 1 day: the log. Record the decision, the questions you did not predict, and the tripwires you now report against. (10 minutes)
Total cost: roughly four hours per high-stakes recommendation. The comparison price is one meeting where a stakeholder finds the flaw you did not.
Scaling the protocol to the stakes
Not every recommendation needs the full run. A rough scale:
| Stakes | Run |
|---|---|
| Routine budget item, reversible | Seven-pattern scan and a tripwire. 20 minutes. |
| Significant investment or policy change | Full stress test, question bank, one-pager. 2.5 hours. |
| Board-level, irreversible, or reputationally exposed | Full protocol including murder board and pre-read check. 4 hours. |
The scan is never skipped. It is the twenty minutes that catches the hockey stick before it reaches anyone.
Institutionalizing the method
The fastest way to raise the quality of what reaches your desk is to make the stress test the visible standard for getting there.
- Require the claim stack. Any recommendation above a threshold comes with its claims and assumptions listed. The format does the work; people write differently when they know the assumptions will be visible.
- Require the alternatives. The do-nothing case and the staged case, with numbers, in every significant proposal.
- Require the tripwires. No approval without kill criteria. Then report against them.
- Run the personas in review. Before a recommendation goes up, someone plays the verifier and someone plays the operator. Twenty minutes, in prep, not in the meeting.
- Use the checklist to improve decisions, not to ambush colleagues. "Have we tested the do-nothing case?" lands better in prep than in the meeting. The goal is a team whose recommendations arrive stress-marked, not a culture where presenting is dangerous.
What you built in this course
The next 30 days
Days 1 to 7: Finish the full stress test on your real case. Steps 3 through 5, then the rebuild. Use the block you scheduled in Module 3. Complete the one-pager.
Days 8 to 14: Complete the preparation cycle. Question bank, answers, pre-read check, murder board. If your real presentation lands in this window, run the full T-minus protocol against it.
Days 15 to 21: Run the seven-pattern checklist on incoming recommendations. The patterns are as common in what you receive as in what you produce, and spotting them in others' work sharpens your eye for your own. Use it in prep, not in the meeting.
Days 22 to 30: Institutionalize. Add the claim stack, the alternatives requirement, and the one-pager to how your team prepares recommendations for you. Report against your own tripwires at the first cycle after approval.
The one-month test
After your next high-stakes presentation, score it on one measure: how many questions in the room were already in your bank?
Above 80 percent means the protocol is working; the meeting held no surprises because you had already run it. Below that, compare the missed questions against the personas and find which lens you under-weighted. The bank improves every cycle, which is the compounding at the center of this course: each defended recommendation makes the next one cheaper to defend.
Final principle
Everything here reduces to one sentence: AI made scrutiny fast and cheap, so the leaders who thrive are the ones who point that scrutiny at their own thinking first.
Your stakeholders' tools will keep getting sharper. So will yours. The advantage never belonged to whoever had better tools. It belongs to whoever is more willing to find out they are wrong while it is still cheap to be.
Appendix A: Quick Reference
The seven weaknesses
Buried assumption. Chosen baseline. Hockey stick. Friendly pilot. Missing alternative. Convenient source. All-or-nothing ask.
The detection questions
What else has to be true? Compared to what, chosen by whom? What produces the bend, and when is it visible? How does the pilot differ from the rollout population? Where does the strongest alternative lose? Who produced it, wanting what, when? What would make us stop, and when?
The stress test
Strip to the claim stack. Rate assumptions (load-bearing × uncertain). Attack the evidence (source, age, denominator, counterfactual, survivors, reference class). Compare against steelman alternatives. Break it with a pre-mortem, then set tripwires and responses. Rebuild.
The personas
Verifier. Capital allocator. Operator. Historian. Plus regulator and affected party where the room needs them.
The answer architecture
Direct answer first sentence. Evidence second. Implication for the decision third. Also: the concession, the reframe (sparingly), the bridge.
The can't-answer protocol
Never improvise a number. Bound what you do know. Commit to a date, then deliver.
The one-pager
Recommendation. The case. Load-bearing assumption. Alternatives considered. Risks and responses. What would change my mind.
The pre-presentation protocol
T-7: stress test and rebuild. T-4: question bank and answers. T-3: pre-read check. T-2: murder board. T-1: one-pager and disclosure. Day of: nothing new. T+1: log it.
Key evidence
PwC survey of corporate directors: 35 percent say their boards have integrated AI into oversight activities; PwC's companion survey of C-suite executives found 99 percent believe boards should be using AI in oversight. Prospective hindsight research: imagining an outcome as already having occurred, then explaining it, materially increases the number and specificity of causes identified. Planning fallacy and reference class forecasting: inside-view projections systematically run optimistic; anchoring on how similar initiatives actually performed corrects for it.
Appendix B: AI Toolkit
Copy-Paste Instructions for the Full Protocol
All prompts are tool-agnostic. Replace anything in [brackets]. Confidentiality first: follow your organization's policy on sharing strategy material with AI tools. Every prompt below works on an abstracted version of your case: keep the logic and structure, round the numbers, remove names and identifying details.
The seven-pattern scan
Here is a strategic recommendation: [paste].
Check it against seven patterns and, for each, say
"present," "absent," or "cannot tell," with the
sentence that triggered your judgment:
1. Buried assumption: a conclusion whose required
conditions are not stated
2. Chosen baseline: a comparison whose reference point
flatters the case
3. Hockey stick: a projection that bends away from
history without a named mechanism
4. Friendly pilot: evidence from favorable conditions
generalized to unfavorable ones
5. Missing alternative: no do-nothing or staged case
6. Convenient source: evidence from an interested or
stale source
7. All-or-nothing ask: no gates, no stopping conditionsThe stress test
Step 1: Strip to the claim stack
Here is a strategic recommendation: [paste document or
abstracted version].
Reduce it to a claim stack:
1. The recommendation in one sentence, including the
ask and its structure
2. The 3 to 5 key claims that must be true for the
recommendation to hold
3. Under each claim, every assumption it rests on,
including unstated ones the text takes for granted
Flag any assumption that appears nowhere in the
document but is required by its logic.Step 2: Rate the assumptions
For each assumption in this claim stack, rate:
1. Load-bearing: if this is wrong, does the
recommendation survive? (fatal / damaged / survives)
2. Uncertainty: how strong is the stated evidence?
(strong / thin / none stated)
List the assumptions that are both fatal-if-wrong and
thin-or-no evidence. That is my attack list. Rank it
by how easily an outside stakeholder could challenge
each one with public information.Step 3: Attack the evidence
Here is the evidence supporting a load-bearing claim:
[paste evidence summary].
Interrogate it adversarially:
1. Source: who produced each item, and what did they
want to be true?
2. Age: when is each from, and what could have changed?
3. Denominator: what population does it describe, and
is that the population being projected onto?
4. Counterfactual: compared to what? What baseline
choice is doing silent work?
5. Survivors: does it look only at winners?
6. Reference class: for any projection, what class of
similar initiatives should it be anchored to, and
what does that class typically deliver?
Do not soften. For each vulnerability, write the exact
question a hostile stakeholder would ask.Step 4: Steelman the alternatives
The recommendation is: [one sentence].
Build the strongest honest case for:
1. Doing nothing, including everything the money,
time, and attention could do instead
2. A staged or smaller reversible version
3. The strongest genuinely different approach
Argue each as its best advocate would, not as a straw
man. Then state what evidence would have to be true
for each alternative to beat my recommendation, and
estimate in numbers where the recommendation beats
each one.Step 5: Break it
It is [18 months] from now and this initiative failed.
Write the 3 most plausible failure stories, each with:
1. The specific chain of causes
2. The earliest observable evidence that this failure
was beginning (the tripwire), with a number and a
date where possible
3. What the reasonable response would have been at
that tripwire moment
Prioritize failure modes that stem from the
assumptions in my attack list.The rebuild check
Here is my revised recommendation: [paste].
Verify it now contains: assumptions stated with the
load-bearing one flagged; projections anchored to a
reference class; alternatives shown with the reason
each lost; risks with tripwires and responses; and a
defined condition under which the initiative stops.
List anything still missing, and identify the weakest
remaining claim.The question bank: personas
Run each persona separately against your rebuilt document, then merge.
The verifier
You are a board member who independently checks
management's numbers. You have this document, public
benchmark data, and the presenter's track record:
[past initiatives and outcomes, abstracted].
Generate your 8 hardest questions. Prioritize places
where the document's numbers can be checked against
external data or the presenter's own history.The capital allocator
You are a director who treats every ask as competing
with every other use of capital. Generate your 8
hardest questions about this recommendation,
prioritizing opportunity cost, the alternatives not
shown, and whether the ask's structure (size, timing,
staging) is justified.The operator
You are an executive who has run implementations this
size. Ignore the strategy; attack the execution.
Generate your 8 hardest questions about capacity,
sequencing, hiring, timelines, dependencies, and what
this initiative does to the performance of existing
operations while it absorbs attention.The historian
You are the longest-tenured person in the room. You
remember every similar initiative: [list prior
comparable efforts and outcomes, abstracted].
Generate your 8 hardest questions connecting this
recommendation to that history, especially "what did
we learn last time and where does this plan apply it?"The regulator (where relevant)
You are the person accountable for compliance, safety,
or licensure exposure. Generate your 6 hardest
questions about what this initiative risks on that
front, who signs off before each stage, and what
happens to our regulatory standing if it goes wrong.The affected party (where relevant)
You represent the staff, clients, or customers who
will live with this initiative. Generate your 6
hardest questions about what it costs them, whether
they were consulted, and what happens to them if the
tripwires trip.Merge and rank
Here are the question lists: [paste all].
Deduplicate, then rank the top 20 by how much damage
an unprepared answer would do to the recommendation's
credibility. Mark the 3 where you predict I currently
have no strong answer.Answers and rehearsal
The answer architect
Question: [paste one hard question].
My raw material: [facts, numbers, reasoning].
Draft an answer in exactly three parts: the direct
answer in one sentence, the strongest evidence in one
or two sentences, and the implication for the decision
in one sentence. No preamble, no "great question."The answer scorer
Here is a question and my answer: [paste both].
Score the answer: Did the first sentence answer
directly? Is the evidence specific and checkable? Does
it state the implication for the decision? Does any
part read as evasion? Rewrite it tighter if it fails
any check.The murder board
Run a live hostile Q&A on this recommendation: [paste
one-pager or rebuilt document].
Rotate through four personas: a verifier who checks
numbers, a capital allocator focused on opportunity
cost, an operator attacking execution, and a historian
citing past initiatives: [abstracted history].
Ask one question at a time and wait for my answer.
After each answer, follow up once the way a skeptical
stakeholder would: probe the weakest part of what I
said. Every 5 questions, pause and score my answers
against this standard: direct answer first, evidence
second, implication third. Do not be polite. Begin.The bounded non-answer
I will likely be asked: [the question I cannot fully
answer]. What I do know: [adjacent facts, bounds,
trends]. Draft a response that: states plainly what I
do not have, bounds the uncertainty with what I do
know, and commits to a specific delivery date. No
bluffing, no filler.The pre-read and the one-pager
The skeptical reader (run against your pre-read)
You are a skeptical board member who received this
pre-read two days before the meeting: [paste].
Generate the 20 hardest questions it raises, with the
section each comes from. Then mark which of the 20 the
document already answers adequately, and which it
does not.The one-pager
From this stress-tested recommendation: [paste rebuilt
document], generate a one-page decision summary with
exactly these sections:
1. Recommendation (one sentence, including the ask
and its structure)
2. The case (3 to 5 claims, one line each)
3. Load-bearing assumption (named, with its evidence
and its tripwire)
4. Alternatives considered (each with the one-line
reason it lost)
5. Risks and responses (top 3, each with a tripwire)
6. What would change my mind
Keep it under one page. Where my document lacks the
content for a section, write [NEEDED: description]
rather than inventing it.The proactive disclosure
Here is my one-pager: [paste]. Identify the single
weakest element a hostile reader would find first.
Draft one sentence, for me to say early in the
meeting, that names that weakness plainly and states
what the plan does about it.The tripwire report
Here are my approved tripwires: [paste] and the
current numbers: [paste]. Produce a status line for
each (green / watch / tripped, with the number) and,
for any that is tripped, restate the pre-agreed
response as this cycle's recommendation.Standing instructions worth saving
Add to your AI tool's custom instructions, or to preferences.md if you built a second brain:
When reviewing my strategic recommendations or
decision documents:
- Act as a red team by default; do not validate
- Surface unstated assumptions before commenting on
stated ones
- Always ask "compared to what?" about every
favorable number
- Flag any projection that bends away from historical
trend without a named mechanism, and ask for the
reference class
- Never invent facts or figures; mark gaps as
[NEEDED: description]
- When I ask for a defense of my position, first give
me the strongest attack on it
- When I present a full commitment, ask what the
staged version would look like