---
title: The Public Data Layer
slug: public-data-layer
works_with: [Google Sheets, Clay, HubSpot, Airtable]
type: template
audience: RevOps, SDRs and AEs doing their own research, founders running outbound before hiring for it, GTM engineers building a scoring pipeline
promise: Score and sort a list of target accounts into act, watch, or drop, using only free public sources, before a single paid enrichment credit gets spent.
time: 45 min
format: markdown + CSV template
cta: If a segment keeps landing at three flags no matter which free source you try, that's a Free Symptom Read conversation, not a bigger spreadsheet.
category: account research and prioritisation
pairs_with:
  - slug: observed-vs-inferred-outreach
    why: turns the evidence this sheet produces into a line you can actually send
  - slug: symptom-read
    why: bring the scored sheet to the call instead of a cold guess
  - slug: scientific-protocol
    why: the abstention and convergence rules here are the same discipline the protocol runs on every claim
---

This is the layer you run before you pay for anything. Most account-research tools sell the enrichment step and skip the two steps that make enrichment worth buying: a disciplined order for where you look, and a rule for what happens when the answer is genuinely unknown. Both are free. Neither takes more than an evening for a first batch.

The output isn't a lead list. It's a sheet with three live states — act, watch, drop — and a paper trail behind every point on the score. If you can't point to the URL and the date behind a number, the number doesn't go in the sheet.

## Free sources answer most of this before a paid one gets a turn

Work this list top to bottom. Each row is free — no key, no card, no trial clock.

| Source | What it proves | How to pull it | Signal weight |
|---|---|---|---|
| Company careers page + ATS JSON | A dated, provable open role | Greenhouse: `boards-api.greenhouse.io/v1/boards/{token}/jobs?content=true` → `first_published`. Lever: `api.lever.co/v0/postings/{slug}?mode=json` → `createdAt`. Ashby: `api.ashbyhq.com/posting-api/job-board/{boardName}` → `publishedAt`. Paste the URL in a browser tab — it's plain JSON, no auth. | Primary |
| SEC EDGAR full-text search | A recent funding round (Form D) | `efts.sec.gov` full-text search, or the EDGAR company search, by company name | Corroborator |
| Wayback Machine (web.archive.org) | A date on a page that shows no timestamp itself | Paste the exact job or pricing-page URL, compare snapshot dates. If the same posting shows up across 45+ days of snapshots, that's dated, citable proof it's been open that long. | Dating tool for any row above |
| LinkedIn People tab, checked by hand | Someone left a relevant seat, or a new leader started | A human-run, curated check on a short account list — not a scraper, not an automated alert | Corroborator / timing |
| G2 / Capterra reviews (dated) | A tool was bought but never fully rolled out | Search company + tool name, read the dated review text for "inherited," "messy," "never fully implemented" | Corroborator |
| Company website, pricing page, demo form | A funnel that can't route its own inbound | Visit it yourself | Weakest — season with this, never lead on it |
| Google Alerts / RSS on the company name | Leadership hires, press, funding announcements | Free alert, arrives on its own | Corroborator / timing |

Climb past this list only when a specific record clears the score below and still has no name attached to it — that's the one job worth paying for. Apollo or Hunter for a verified email, Clearbit or ZoomInfo for firmographic depth, in that order, and only after the free pass. Cache whatever you buy by domain so the same account never gets paid for twice.

## Every account gets the same columns, win or lose

| Field | What goes in it | Where it comes from |
|---|---|---|
| `domain` | The dedupe key — one row per domain, always | The company's own site |
| `company_name` | Display name | Site or filing |
| `signal_type` | `open_req` / `funding` / `personnel_departure` / `leadership_change` / `tool_review` / `website` | Matches the source table above |
| `source_class` | `ats_api` / `sec_edgar` / `wayback` / `linkedin_manual` / `review_site` / `press` / `own_site` | Which row of the source table fired |
| `signal_date` | The real date behind the signal — never today's date | ATS `first_published`/`createdAt`/`publishedAt`, filing date, or Wayback snapshot date |
| `source_url` | The exact link | Copy it, don't summarize it |
| `evidence_quote` | The exact sentence you're relying on, pasted verbatim | The source page itself |
| `company_size_band`, `department`, `seniority`, `business_model`, `geography_tier` | The five firmographic facts the rubric below scores | Site, LinkedIn, filing |

## One hundred points, six factors, and a column for what you don't know

Score every account the same way, in the same order. Each factor states what earns points and — separately — what happens when you genuinely cannot find the answer. "Genuinely cannot find" is not the same as "found and it's low." A known junior title scores its stated points, not a flag.

| Factor | Points | If you have the evidence | If it's genuinely unknown |
|---|---|---|---|
| Company size | 40 | 1,000+ = 40 · 200–999 = 30 · 50–199 = 20 · 10–49 = 10 | 0 + flag |
| Department match | 15 | Target department = 15 · adjacent = 7 | 0 + flag |
| Seniority | 15 | VP+ = 15 · Director = 10 · Manager = 5 · known IC = 0 | 0 + flag only if no name or title can be found at all |
| Business model | 10 | B2B = 10 · B2C = 0 | 0 + flag only if it's genuinely indeterminate |
| Geography | 10 | Tier-1 market = 10 · Tier-2 = 5 · known Tier-3 = 0 | 0 + flag only if the company can't be located at all |
| Signal type | 10 | Primary dated signal (open req, funding, confirmed departure) = 10 · secondary/corroborating mention only = 6 · pure firmographic prospecting, no specific trigger = 2 | 0 + flag if nothing at all was found |

**Thresholds:**

- **≥ 70 → Active.** Enough here to spend an hour writing to this account.
- **45–69 → Watch.** Re-check in 2–4 weeks — the top-weighted factors (size, req age) are the ones most likely to still move.
- **< 45 → Drop.** Doesn't clear the bar even generously.
- **Any record with 3 or more "0 + flag" cells → Abstain**, regardless of what the point total says. See below.

## An unknown scores zero, never the average

It's tempting to split the difference on a field you're not sure about — write 20 instead of 40, or 7 instead of 15, so the sheet feels less empty. Don't. An unknown that scores an average is a guess wearing a number, and every stage downstream treats it as if it were evidence.

Score it zero. Mark which factor was unscoreable in the flag column. Zero is honest about what you don't know; an average hides it. BoldGTM runs this same rule on its own scoring work and calls it the Abstention Rule — a record doesn't earn partial credit for a fact nobody could verify.

## The record needs two kinds of source before it moves, and three flags ends it

Run this once a week, before anything moves from Watch or Active into outreach prep.

**Check 1 — corroboration.** Count distinct `source_class` values on the record, not distinct mentions.
- One source class = **LOW** confidence. Don't act on it yet — one loud source is one source repeated.
- Two distinct classes = **MED** confidence. Enough to draft.
- Three or more distinct classes = **HIGH** confidence. Enough to lead a message with it.

**Tie-break when two sources back one row.** The row's `source_class` is the strongest one by
the source table's Signal weight (Primary beats Corroborator beats Dating tool beats Weakest).
The others don't disappear: paste each extra source's URL and quote into `evidence_quote`,
one per line — they are what the corroboration count is made of.

This is the same logic BoldGTM's site calls the Convergence Gate: three citations of the same kind aren't three citations, they're an echo.

**Check 2 — the flag count.** Add up the "0 + flag" cells from the rubric. Three flags on one record, no matter what the other factors add up to, means the record leaves the funnel. Move it to an Archive/Unscoreable tab — not the trash. New evidence can bring it back later. A 90-point record with three unknowns is three unknowns wearing a 90; don't let the other factors talk you out of the exit.

## Sheet: tab `accounts`, one row per domain — paste this into row one

Two tabs. `accounts` holds every scored record; `archive` is where a record goes at three
flags — same columns, never the trash.

```csv
domain,company_name,signal_type,source_class,signal_date,source_url,evidence_quote,company_size_pts,department_pts,seniority_pts,business_model_pts,geography_pts,signal_type_pts,total_score,unknown_flags,confidence,tier,status,next_action,reviewed_by,review_date
```

One filled row, composite example — not a real company:

```csv
acmedata.io,Acme Data (composite),open_req,ats_api,2026-06-02,https://boards.greenhouse.io/acmedata/jobs/1234,"RevOps Manager — Remote, posted 2026-06-02",30,15,10,10,10,10,85,0,HIGH,Active,ready-for-review,draft outreach,,
```

**Legend:**
- `tier` ∈ {Active, Watch, Drop, Abstain}
- `status` ∈ {sourcing, scored, ready-for-review, ready-for-outreach, archived}
- `confidence` ∈ {LOW, MED, HIGH} — set by the corroboration check, not the point total

## Pairs with

Run [Observed vs Inferred Outreach](observed-vs-inferred-outreach.md) right after this one — the "observed" column in that worksheet should come straight from the `evidence_quote` and `source_url` fields this sheet already dated and sourced, not from memory.

If a batch clears the rubric and the confidence check but still won't turn into a sendable line, bring the sheet to [Free Symptom Read](symptom-read.md) instead of reworking it alone — that's a fifteen-minute conversation, not a bigger spreadsheet.

The abstention rule and the corroboration check here aren't house style — they're the same two rules the [Scientific Protocol](scientific-protocol.md) runs on every claim BoldGTM makes about a client's system. Worth reading if the discipline in this sheet is the part you actually came for.
