← Free GTM Hacks

The BoldGTM Scientific Protocol

Twelve rules for telling a real GTM cause from a plausible story, so you stop paying to fix the wrong thing.

Works in
Google Sheets · Airtable · Notion
Format
markdown protocol
Time
20 min read
For
founders, GTM operators, anyone about to spend budget or headcount on a GTM fix
Job
GTM diagnostic discipline

Twelve rules for not fooling yourself while diagnosing a GTM problem. The weighting formula and the evidence thresholds below are stated exactly, so a person or an agent applies them the same way every time. Nothing here is automated for you. It's a discipline you apply by hand, the same way every time, whether or not anyone is checking.

Nothing here is certified or audited. It's a working method, versioned, applied to BoldGTM's own diagnoses before it's handed to anyone else's.

1. Never treat a symptom as a problem.

A reported pain — "outbound isn't working," "reps keep asking the same thing" — is a symptom observation. It is not a problem statement. The problem is whatever survives after diagnosis, and it is usually not the sentence you started with.

Prevents: re-fixing the same thing under a new name every quarter. Test: name what changed the last two times you "fixed" this. If the fix was different each time but the symptom came back identical, you were patching a symptom, not solving a problem.

2. Never treat an interpretation as an observation.

"The team is slow" is an interpretation. "Deal 14 sat eleven days between demo and follow-up" is an observation. Keep them in different columns — the moment they merge, your notes stop being evidence and start being a story you already believe.

Prevents: building a fix on a mood instead of a fact. Test: read your last written note on this problem. Does it contain a date, a count, or a name — or a word like "always," "never," "seems," "slow"? The second kind is an interpretation wearing an observation's clothes.

3. Never treat correlation as mechanism.

Two things moving together supports a hypothesis. It never settles one. A claim resting only on co-occurrence cannot outrank "moderate" confidence, no matter how clean the chart looks.

Prevents: paying to fix the thing that happened to move at the same time as the symptom, instead of the thing that caused it. Test: finish the sentence "X causes Y because ___." If you can't fill in the mechanism, you have a correlation, and it caps your confidence — it doesn't end the inquiry.

4. Define the construct before you measure it.

Before you can count "engaged" accounts or "qualified" leads, the word needs an operational definition with named indicators. An ambiguous construct is a moment to stop and define, not a number to start tracking.

Prevents: a metric that moves for reasons nobody can explain, because it was never pinned down in the first place. Test: write the one-sentence definition of the word you're using. If you can't, you don't have a metric — you have a mood with a dashboard.

5. Hold at least two live explanations, not one.

One explanation is a bet, not an inquiry. The moment you commit to a single story, every new fact gets read as support for it — that's not analysis, that's advocacy.

Prevents: confirmation bias dressed up as a finding. Test: write a second explanation for the same symptom, one you find at least somewhat plausible. If you can't, you haven't looked yet. You've decided.

6. For every explanation, write what would kill it — before what would prove it.

State the disconfirming evidence first. An explanation that can't say what would reject it isn't a hypothesis. It's a conclusion wearing a hypothesis's clothes.

Prevents: the single most expensive habit in diagnosis — looking only for what agrees with you. Test: for your leading explanation, write the one observation that would make you drop it. If nothing comes to mind, you haven't made it falsifiable, and building on it is a bet.

7. Separate what you observed from what you concluded — in writing.

Every statement that goes beyond the raw record is an inference, and it should read like one. "Three reps stopped logging notes in week two" is a citation. "The tool is too slow for reps to bother" is an inference sitting on top of it.

Prevents: notes that blur into a narrative until nobody can tell, six weeks later, what was actually seen versus assumed. Test: mark your notes [O] for observed, [I] for inferred. A document that's mostly [I] with no [O] underneath isn't evidence. It's an opinion with formatting.

8. Grade support by independent evidence classes, not by count.

Three interviews saying the same thing feel like triangulation. They're one witness, repeated three times. Support is graded by distinct evidence classes — interview, CRM record, financial, complaint log, direct observation, experiment — not by how many rows you collected.

Prevents: an echo mistaken for corroboration. Test: list your evidence, then list the class each item belongs to. BoldGTM's own Convergence Gate calls a claim HIGH only at ≥3 sources and ≥2 distinct classes and mean quality ≥0.6; same-class sources, however many, cap out at LOW.

9. Report confidence as a grade, not a number you can't defend.

Confidence renders as weak, moderate, or strong — never a decimal, never a percentage — until enough settled outcomes exist to calibrate what those numbers would even mean.

Prevents: false precision: "73% confident" when nobody has ever checked what 73% correlates with. Test: can you state why this is weak, moderate, or strong in one sentence citing your evidence classes? If the only justification is a feeling, it's a guess with a costume, not a grade.

10. Recommend the cheapest experiment that would actually move your uncertainty — kill criterion first.

When evidence stalls, the next step is a test, sized by cost and chosen for the information it would produce — and the number that would kill it gets written down before the test runs, not after the results are in.

Prevents: expensive fixes (a hire, a rebuild) launched on evidence a free test could have gathered first — and tests that quietly never get called a failure because nobody agreed in advance what failure looked like. Test: before you start, write the number that would make you stop. If you can't write it before you have data, you'll rationalize whatever number you get after.

11. Update the grade the moment contradicting evidence appears — visibly.

A belief that changes quietly in your head but never changes what you tell your team isn't an update. It's private doubt sitting next to a public plan that no longer matches it.

Prevents: an organization acting on a diagnosis the diagnoser has already abandoned. Test: find the last time you were wrong about a GTM cause. Is there a dated note anywhere that says so, or did the old explanation just quietly stop being mentioned?

12. Abstain when the evidence can't support a conclusion — and say so.

When the Convergence Gate isn't met, the honest output is not a downgraded recommendation. It's a stated abstention: what's missing, and the cheapest way to get it. Abstention is an answer, not a failure to produce one.

Prevents: the recommendation everyone actually makes when they're stuck — guessing, and calling it a decision. Test: count your genuine unknowns on the questions that matter — not "haven't thought about," but no dated fact or named instance behind them. Three or more, and the record should leave the funnel rather than enter it on a guess.

Sheet: two tabs, claims and evidence. The gate reads them, it doesn't guess.

A claim is one row. Its evidence is many rows, each with a class and a quality weight. The Convergence Gate (rule 8) is computed from the evidence rows, never typed by hand: n = distinct source rows, d = distinct classes, q = mean quality.

claim_id,claim,pillar,owner,n,d,q,grade,confidence_word,updated,what_would_change_it
C-001,"Feedback has no owner",sales enablement,"(name)",2,2,0.9,MED,moderate,2026-09-06,"A third source of a new class: support tickets or churn-survey comments"
claim_id,class,quality,source,date,observed_or_inferred,quote
C-001,interview,1.0,"Rep interview #1",2026-08-28,O,"We answer the pricing objection live on every call"
C-001,crm_record,0.8,"HubSpot export, deals lost Q3",2026-09-01,O,"Lost reason: pricing question unanswered (14 of 61)"

class ∈ {interview, crm_record, financial, complaint_log, direct_observation, experiment}. quality: first-party 1.0 · corroborated 0.8 · single-operator 0.6 · mixed 0.5. grade: HIGH at n ≥ 3 and d ≥ 2 and q ≥ 0.6 · MED at n ≥ 2 and q ≥ 0.5 · else LOW · ABSTAIN when the gate isn't met and you say so (rule 12). confidence_word ∈ {weak, moderate, strong} — never a decimal (rule 9). Composite rows; not a client.

A rule you can't apply alone is a decision, not a rule.

Rule 4 and rule 8 both have a real failure mode: sometimes the construct is ambiguous, or two evidence classes flatly disagree, and no amount of staring at your own notes resolves it. That's not a sign the protocol failed — it's the specific moment a second, disinterested reader is worth more than another hour alone with the data. That's what the actual conversation is for.

Step 1 of 3

Where does the symptom keep returning?

Free 30 minutes. Bring the recurring symptom; we will test whether problem mapping is the right next move.