GTM,
engineered

Shamas Hamad · GTM Engineer

I find the moment a company becomes a buyer, in public data, then build the machine that acts on it.

One of five speakers at a Claude Code conference in London during London Tech Week, on pattern interruption in go-to-market · 20+ systems across 13+ clients in 6 countries · BSc Computer Science, data science specialism

01 · The interview offer

Give me an interview and I turn up with your pipeline already drawn.

Two things, built from public sources before we speak: a tailored 30/60/90 for the actual role, and an Excalidraw map of your pipeline from signal to signed deal. Both are wrong in places on purpose, because the fastest way to learn a company is to be corrected on something specific.

  • The 30/60/90. What I would baseline in week one, which of your signals I would back-test against deals you already won, what I would ship small and gated before day 30, and the stop-gate that would make me kill each play rather than scale it. It states its own limits. Where I am guessing at your stack, it says so.
  • The map. Public triggers on the left, each with the actual dataset underneath it, fanning into scoring and fit, then outreach, then the meeting, then the loop that feeds what you learn back into targeting. One page. Editable. Yours to redraw on the call.

The diagram, from a real one · built before a real interview

GTM top-of-funnel diagram: five public triggers, each with its data source, fanning into fit scoring, outreach, and a learning loop Open full size ↗
Where this came from, and why the company name is not on it

Built for a capital-advisory firm ahead of a GTM Engineer interview, from their published sector list and public site copy only. Every trigger on it maps to a named public dataset. Nothing on it is confidential.

Source: the diagram is generated from a checked-in Excalidraw scene and a written content spec, both in my repo. Redaction note: the company name has been removed from the diagram and from every repo path on this page, because that process may still be live. The thinking is the point; the client is not. Paths appear here as job-hunt/[redacted]-interview/….

02 · The first 90 days

The shape every one of those roadmaps takes, generalised from two I have written.

Learn what a good deal looked like before it arrived, make what survives repeatable, then let the unit economics decide what continues.

  1. 1–30

    Find out what a good deal looked like before it arrived.

    The moves
    • Baseline the funnel. Where deals come from today, referral versus sourced, what converts, what stalls. Numbers first, opinions after.
    • Back-test signals against your own won deals. For each past win: was there a public trigger visible before they arrived, and how early?
    • Build the entity spine. Filing entity to company to domain to person to contact, each hop carrying source, date and confidence.
    • Run one small approved cohort, small enough that I can explain every send, capturing why the team accepts, rewrites or rejects each one. The rewrites are the real training data.
    • Fix one contained reliability problem found along the way.

    Ends with: a funnel baseline, a back-test naming which triggers actually preceded good deals, a working entity spine, one approved cohort, one reliability fix.

  2. 31–60

    Make what survived repeatable.

    The moves
    • Enrichment with provenance. Every field carries source, date and confidence, so no message rests on a fact nobody can trace.
    • Suppression as architecture. Customers, live opportunities, partner and investor overlaps, prior passes inside their cooling window, opt-outs. Suppression runs before approval, not after.
    • Signal decay. Triggers get time-weighted and expire, so the pipeline does not silently fill with stale reasons to call.
    • Attribute replies back to the trigger that produced the send, not to the subject line.
    • Multi-thread the economic buyer and the role whose arrival was the trigger.

    Ends with: two or three signal families running as a monitored loop, provenance and suppression enforced in the pipeline, and a written record of which hypotheses were kept and which were killed.

  3. 61–90

    Economics, decay, and what compounds.

    The moves
    • Cost per qualified conversation, by signal family. By day 90 evidence should be deciding which families continue, not preference.
    • Retire and replace as routine. Signal families decay as competitors find them, so a standing backlog makes retiring a weak one a normal Tuesday.
    • Re-score the archive. Everything screened and passed over stays in the pool with the reason attached. When criteria shift, yesterday's near-miss becomes today's first call at no new sourcing cost.
    • Decide on paid, and only if an organic angle proved itself. Narrow air cover onto the matched list, not broad spend.
    • Hand over the next-quarter recommendation: keep, change, stop, with the evidence and the known uncertainties stated plainly.

    Ends with: signal families ranked by cost per qualified conversation, a documented sourcing spine with runbooks, a live experiment backlog, and a plan that says what to stop as clearly as what to scale.

A verbatim slice of a real 30/60/90

The move most signal projects get backwards. "Take past engagements and ask a single question of each: was there a public trigger visible before they arrived, and how early? This inverts the usual order. Most signal projects pick plausible triggers and hope; I would rather learn which triggers actually preceded deals [they] already value."

The unglamorous part that decides whether any of it works. "A filing entity name rarely matches the operating brand, which is where most signal work quietly fails. Filing to company to domain to founder to contact, each hop carrying source, date and confidence. That is the difference between a signal list and a usable pipeline."

What I refuse to promise. "Baselines are agreed after observing the current system, not invented in advance. Nothing in this plan promises a number I have no basis to promise."

Source: job-hunt/[redacted]-interview/30-60-90-content.md, lines 45, 47, 112, company specifics removed. A second, independently built roadmap for a different company exists at job-hunt/interviews/[redacted]/30-60-90-seed.md, which is how I know the method travels rather than being one lucky document.

03 · The numbers

Outbound numbers from delivered client work, in the client's own reporting.

Outbound is the system with the longest run behind it, so these are its numbers. Paid ads, content, AI search and client management have their own.

Status: client-reported and client-credited figures from delivered engagements. Measured outcomes, not projections, and not the packaged system I sell today. Why they are here: an agency or founder background usually reads to a hiring manager as execution without infrastructure. This is the execution half. The next four sections are the infrastructure half, which is the part that is harder to fake. Full context on how each was measured is on the main portfolio, or ask me and I will walk you through any of them.

04 · GTM systems · the signal layer

A signal is the moment information arrives, not the state it describes.

The next four sections are one system, in four layers: the signal spec, the engine, the experiment loop, and the call loop. It runs on my own pipeline every weekday. Not a portfolio piece, which is why it has monitoring, audit logs, a post-mortem, and a gate that has told me no.

Everyone says "signal-based". Most of what gets called a signal is a description of a company, which is useless, because it was true last year and will be true next year. The rule above does most of the work. The second rule does the rest: every signal has to answer "why does this company need what I sell, now?"

TriggerPublicly observable inWhy it means "now"
Mid-raise, round not filledSEC Form D, item 13.cItem 13.c is literally "Total Remaining to be Sold". A dated public record of a company actively raising and short of target.
16 months post-seed, no Series Around-date dataMedian seed to Series A gap is around 20 months. Past that point, the clock is the pressure.
First-ever CFO or Head of Finance postedjob boards, LinkedInThe first finance hire is a company admitting money has become the constraint, usually before it says so anywhere else.
Contract award larger than anything delivered beforeSAM.gov, USAspending, Find a Tender, TEDDelivery has to be funded, planned or not. Comparing an award against that same supplier's own prior awards needs no balance sheet, which is why it works across jurisdictions.
Clinical phase moves upClinicalTrials.govA phase transition is a dated, public commitment to a materially larger spend.

Source: the signal set from the diagram in section 01, sources as drawn. Median seed-to-Series-A figure cited in my spec to Carta Q2 2025 (616 days). These are the triggers for one specific buyer. The transferable part is the shape: an event, a named dataset, and a distinct reason for urgency.

Why those two rules, and what they cost

On rule 1. "Company is profitable" is a trait. "Company filed accounts showing a third profitable year" is a signal, because the filing happened on a date and is publicly observable. Traits give you a list. Only events give you a reason to call this week rather than any other week.

On rule 2. If four triggers all mean the same thing, that is one trigger with four sources. A good signal set is a set of genuinely different reasons the same problem becomes urgent. That is the difference I look for when someone shows me their targeting.

Source: both rules are written as governing constraints in my own diagram content spec, under "Governing rule for signals". They exist because an earlier version of that diagram was rejected for putting a trait on it, and I wrote the rule so it could not happen again.

Two guardrails I enforce in code, not in a document

A signal I cannot point at is a guess with better branding. In my own system a signal scored above 7 has to carry a source URL or the record fails validation and never reaches a message. It is a schema rule, checked by code before anything is written back.

Some signals point the wrong way. "I treat hiring as an anti-signal in my own lane. A company hiring a GTM Engineer is not a buyer, it is a competitor forming." Knowing which triggers to invert is the same skill as knowing which to chase, and it is cheaper to learn once than to discover across three months of sends.

Source: both stated verbatim in job-hunt/gtm-vocabulary-crib-sheet.md; the source-URL rule is enforced by outbound/auto-outbound/executions/research_schema_validate.py, which is deliberately code-only and does not judge content quality.

05 · GTM systems · the engine

auto-outbound: a stranger to a booked call, with the call data feeding back into the front of it.

73 Python scripts and 24,327 lines across 8 scheduled jobs, 304 commits since December 2025, and 3,156 leads sourced and evaluated all time. Three stages, each one gated.

  1. 1

    Source and qualify

    Scrapes and imports from several lanes, then a two-pass ICP filter: deterministic pre-filters first (geography, freelancer patterns, competitor names, script-ratio checks) to strip obvious rejects before any model call, then a scored model pass. Every source is scored across the funnel separately, and a cohort gate blocks sending for any source under its acceptance floor.

    icp_filter.py · source_scorecard.py · connection_cohort_gate.py

  2. 2

    Research and write

    Each connected lead gets one deep-research pass on a cheap model, validated against a schema by code rather than by another model. The message layer generates variants, selects an image only if a real captured signal exists, and every draft passes a rule-based scanner with 15 blocking rules before it can reach an approval gate.

    research_schema_validate.py · image_candidate_select.py · qa/rules/dm_to_reality_v1.md

  3. 3

    Send, route, recover

    Sends behind explicit flags with per-account daily caps and randomised volume. Replies are classified deterministically into eight classes, and only terminal classes can be auto-dispositioned. A state-legality rule pack blocks illegal transitions, such as advancing a follow-up sequence while a reply sits unhandled.

    unipile.py · classify_reply.py · reply_routing.py · qa/rules/lead_state_consistency_v1.md

Source: file and line counts measured over outbound/auto-outbound/executions/**/*.py; lead count from outbound/auto-outbound/v1-system-map.md line 203, full-universe evaluation dated 2026-07-04; commit count from git rev-list --count HEAD against a first commit of 2025-12-20. Read the authorship disclosure below before reading those line counts as hand-typed.

The failure that taught me to watch the watcher

Live sending stopped for three weeks and nothing told me. An acceptance-rate gate was computing over all time instead of a rolling window, and a partial deploy had shipped the scheduled jobs without the libraries they imported. The pipeline looked calm because a stalled system and an idle system produce the same silence.

What I built after: a deadman monitor that reads each job's own log tail, computes silence against a per-job tolerance, knows the difference between an expected quiet weekend and a real stall, and batches findings into one alert. Every scheduled job now writes a heartbeat even when it does nothing, so "no log line" can only ever mean the cron did not fire.

Source: incident and root cause in outbound/auto-outbound/v1-system-map.md §5b (2026-06-11 to 2026-07-01); monitor in executions/cron_healthcheck.py; heartbeat convention in directives/crontab.md.

A statistics bug I caught in my own gate

The first version of my cohort acceptance gate would have reported close to 100% acceptance for every cohort, forever. Two independent 14-day windows interacted to create survivorship bias: connections that had not yet resolved were being excluded from the denominator in a way that only ever removed the failures.

I found it because the number looked too good, not because anything errored. The explanation, the fix, and a tunable resolution-lag constant are now written into the top of that file so the next person to read it cannot reintroduce it.

Source: executions/connection_cohort_gate.py module docstring and CONNECTION_RESOLUTION_LAG_DAYS.

Integrations, wired rather than listed

Every one of these has real client code in the repo that makes real calls, rather than a logo on a slide: NocoDB as the system of record, Unipile for LinkedIn, Apify for profile and engagement scraping, Tavily and Exa for company-signal discovery and people resolution, Cal.com for bookings, OpenRouter and DeepSeek for qualification and research, Telegram for alerting, ReportLab for generated resource assets, and a self-hosted n8n instance alongside it.

I keep a separate audited file that grades my own tool claims as production, real-but-limited, or name-only, so nothing on a CV outruns what a reviewer would find in the code. I would rather hand you that file than have you find the gap yourself.

Source: integration list per file, in outbound/auto-outbound/executions/. The self-audit is job-hunt/references/claim-boundaries.md, produced by three independent agents reading the code rather than the CV.

Before you run git log: I am not claiming line-by-line authorship

I use Claude Code heavily and my commit history says so plainly. What I am claiming is the architecture, the failure modes, the safety model, and the debugging: the data model, the providers, the gating rules, the schema contracts, the approval policy, the failure handling, and the writeback path.

When the scraper died silently for three weeks, no model found that. When the acceptance gate started reporting numbers that were too good, no model found that either. I found both because I was watching the numbers and knew what they were supposed to look like. That is the job. The typing is not the job.

Source: pre-committed position recorded in job-hunt/references/claim-boundaries.md, "The AI-authorship boundary". The public reference repositories under my GitHub are demonstrations on fictional data and say so in their own READMEs; they are not client work and not this production system.

06 · GTM systems · the experiment layer

"We A/B test our copy" usually means someone eyeballed two numbers at forty sends.

Any prompt surface in the system can register into a shared framework with a real stopping rule. Four constants govern the whole thing.

ConstantValueWhat it governs
Promotion threshold0.85Probability of superiority required before a challenger replaces the baseline, or the baseline kills the challenger
Monte Carlo samples10,000Draws per arm, per read, off each arm's Beta posterior
Volume floor10 per sideBelow it the verdict is "not enough data" rather than a number
Registered surfaces7 (1 live)Declared in YAML with table, reward field and thresholds. A new test needs no new code.

Source: executions/optimization/significance.py, constants DEFAULT_THRESHOLD_PROB = 0.85, DEFAULT_MIN_PER_SIDE = 10, DEFAULT_SAMPLES = 10_000, all verified in file. Surface list and statuses from optimization/surfaces.yaml.

Four design choices I would defend in an interview
  • Arm assignment is a hash, not a lookup. A lead's arm is derived from a hash of the surface name and the lead id, so it is sticky across the whole sequence with no database round trip, and salting the surface in stops two experiments correlating.
  • Adding an experiment is a config row. Surfaces are declared in YAML. A new test needs no new code.
  • The hypothesis generator is forbidden from guessing. It has to find at least three qualifying sources scored on relevance, evidence strength and recency, and if it cannot, it must return "cannot form grounded hypothesis" rather than invent one.
  • The quality scanner must not kill the experiment. A rule pack sitting on top of an A/B test has to distinguish "this is factually wrong" from "this is just the other arm", so variant style differences are explicitly a non-finding.

The maths itself is numpy and about two hundred lines, with a self-test battery of hand-checked cases in the same file. It replaced a fixed-sample heuristic that implied five to six weeks per decision, which at my volume made the loop useless.

optimization/route_variant.py · optimization/registry.py · optimization/surfaces.yaml · optimization/prompts/generator.md · qa/rules/dm_to_reality_v1.md

The hypothesis is written down before the result is read

The current cycle tests message length, cutting the first message from 40 to 60 words down to 25 to 35 with a hard cap at 40. The baseline was recorded first: 45 sends, of which 36 were mature enough to count, producing 6 positive replies, a 16.67% positive-reply rate. The expected lift was declared before any data came back, at a floor of three percentage points and a realistic case of five.

Those lift figures are a stated hypothesis. They are not a result, and I will not present them as one until the posterior says so.

Source: outbound/auto-outbound/optimization/dm1/learnings.md. Baseline figures measured; expected-lift figures explicitly labelled in the file as expected floor and realistic, not outcomes.

07 · GTM systems · the call-data loop

A tool that books meetings never hears the meeting.

Objection data is the highest-quality targeting input available, and it only exists downstream of the send. This is the reason I want the whole funnel rather than the top of it.

  1. 1

    Transcribe locally, for nothing

    Recordings transcribed on my own machine with a local speech model. No per-minute billing, no audio leaving the box, no reason to ration how much gets analysed.

    executions/transcribe_recordings.py (faster-whisper, CPU, offline)

  2. 2

    Match to CRM, flag the inference

    Transcripts matched to CRM rows by content, dispositions pulled from the dialer's API rather than retyped. The review document says out loud that transcripts carry no speaker labels and roles are inferred, and records the one attribution error I made and corrected.

    call reviews/2026-06-30-cold-call-loser-analysis.md, lines 4-6 and 58-60

  3. 3

    Stage the fix behind a volume gate

    Findings become drafted rebuttals in a separate staging file with the evidence quoted next to each one, gated on the pattern repeating across a fuller sample before it can touch the live script.

    call reviews/pending-script-changes.md

The time the call data changed the market, not the wording

The single most common objection across 14 recorded calls, appearing in 3 of them, was a version of "we already do this ourselves". Read as a copy problem, that is a rebuttal to write. Read as a targeting problem, it says something harder: the segment I was calling was capable enough to believe it already owned the thing I was selling.

A month later I cited that exact finding, by file and by count, as the reason a new script dropped that segment from its target market entirely and moved to a different buyer with a different, smaller entry offer. Call data changing the script is table stakes. Call data changing who you call is the version that moves revenue.

Source: objection count from outbound/cold calling/call reviews/2026-06-30-cold-call-loser-analysis.md (3 of 14, measured); the targeting change and its citation of that evidence in outbound/cold calling/playbooks/cold_call_script_v4_signal_feed.md, lines 28-29.

08 · GTM playbooks

Six industries. Each one is a public dataset nobody else is watching.

Same test every time: an event with a date, a named public source, and a distinct reason it means now. Two of these I have run. Four are designed and ready. The mechanics are the claim, and I will defend any of them.

01 Luxury Property Brokerage · Run

Ads as market research, not lead gen

Who follows and engages pool builders and luxury garden designers on LinkedIn.

+

I ran paid Meta ads not to get leads but to learn who the perfect lead is. The data showed our best buyers wanted pools and luxury gardens. So I went where those people already gather: LinkedIn pages of pool construction and garden design firms. I scraped their followers and regular post engagers, filtered them through our ICP, and built a three-month multichannel email and LinkedIn campaign on the result.

Why it works: the ad spend buys a definition of the buyer, not a lead. The definition is then harvestable for free, repeatedly, from a public follower graph.

Path: Meta ad data → LinkedIn company followers and post engagers.

02 Money Lending · Run

Planning permission is a loan application

A freshly approved UK planning application means someone now needs capital to build what they got permission for.

+

I scraped UK planning permission data for approved applications: extensions, conversions, new developments. Every approval is a person or company that just committed to a build and now has to fund it, weeks before they walk into a broker. I matched applicants to Companies House and enrichment data, filtered by project size and type against lending appetite, and built outreach that opens with their own approved project.

Why it works: it reaches borrowers before they shop, which is the only window where a lender is not competing on rate alone.

Path: UK planning permission data → Companies House.

03 Cybersecurity · Designed

The neighbour's house is on fire

A sector peer just appeared in the ICO's public breach register, and every similar firm's board knows it.

+

The ICO publishes every reported data breach, tagged by sector and incident type. Each breached company becomes a flare, not a target. I would harvest its direct peers (same sector, size band, region) through Companies House and Sales Navigator, find whoever owns risk, and open with the fact their insurer will ask about that exact incident at renewal. Three touches, each referencing the real peer breach and the specific control gap.

Why it works: the urgency is external and dated. The prospect did not have to admit a problem for the trigger to be true.

Path: ICO breach register → Companies House, Sales Navigator.

04 Commercial Law · Designed

The losing respondent play

Every published employment tribunal judgment names the employer who just lost, and that firm now knows its legal setup failed.

+

Gov.uk publishes every employment tribunal decision, naming the respondent. A business that just lost or settled has felt the cost of weak contracts or a botched dismissal, in cash. I would pull losing respondents, match them to Companies House for size and directors, filter to commercial firms above a headcount floor, and write one line naming the exact gap the judgment exposed. Then a three-touch sequence to the owner offering the contract and process fix.

Why it works: the judgment is the diagnosis. The message does not have to convince anyone a problem exists.

Path: Gov.uk tribunal decisions → Companies House.

05 EV Charging · Fleet · Designed

The unpartnered salary-sacrifice provider

A leasing or salary-sacrifice EV provider signing up drivers fast, with no charging-payment partner named anywhere on its site.

+

The BVRLA member directory publicly lists every UK leasing and salary-sacrifice provider. I would scrape each member's site and driver FAQ for whether a charging card, home-reimbursement, or roaming partner is named. The ones pushing EV deals hard but silent on charging payment are the gap a fleet-billing platform fills. I would confirm momentum through their EV job ads, then run a three-touch play to the commercial lead, opening with the exact reimbursement gap found on their own site.

Why it works: it is an absence signal. What is missing from a page is as observable as what is on it, and far less contested.

Path: BVRLA member directory → site and FAQ scrape → EV job ads.

06 Recruitment · Designed

The reposted-job distress signal

An employer reposting the same role three or more times has publicly admitted their hiring has failed.

+

I would scrape job boards and career pages daily, keyed on job title plus company. When the same role reappears three or more times over 60 days, that employer is stuck and quietly paying for an empty seat. Filter to the agency's placement niche and seniority, pull the hiring manager and talent lead, and lead the message with their own problem: this role has been open since March, with the days-open count in the first line.

Why it works: the repost count is a measurable, escalating cost the prospect is already feeling and has already published.

Path: job boards and career pages, keyed on title + company, daily.

Status: these are plays, not case studies. "Run" versus "Designed" follows how each is written in my own notes: two are described in past tense as work I did, four in conditional as designs. Deliberately absent: the outcome figures that used to sit on these. They were projections, and a projection printed in past tense reads as a delivered result, so they are gone rather than relabelled. The mechanic is the claim. [UNSOURCED — needs verification] per-play outcome numbers.

09 · Principles

Six beliefs, each one with the decision that cost me something attached.

A principle with no cost attached is a slogan. Open any of them for the decision, the tradeoff, and the file.

01

Live data beats a model of the market, even when the model is bigger

An enrichment model routed roughly 90% of my target market to one offer. Real dials said zero.

+

I had built a routing model over enriched company data that predicted roughly 90% of my agency market should be pitched a client-management system. It was internally consistent, it had volume behind it, and I had already built the script and the collateral for it. Then I made real calls. Not one prospect showed interest in that offer. I reversed the position on a much smaller sample of live conversations than the sample the model was built on, because a dial is ground truth and an enrichment prediction is a hypothesis wearing a percentage.

What it cost: a scripted opener with collateral already built, a rewrite of the live call script, and the reps' practised pitch.

Source: the 90% routing position is recorded in context/records/job-hunt-autoage-dual-track-2026-06-28.md; the reversal, quoted as "live dial evidence (small sample, but zero CM interest so far) beats the enrichment-based ~90% CM-primary routing prediction", is in context/records/future-service-strategy-2026-07-17.md §7a.

02

Instrument the source layer first, because that is where the money quietly leaks

Two lead sources, same ICP filter, same messaging. One accepted at 50%, the other at 16%.

+

I score every lead source across the funnel rather than looking at one blended number. One sourcing lane produced 781 leads at a 50% connection-acceptance rate, a 52% reply rate among those messaged, and a 99% ICP approval rate. Another produced 1,834 leads at 16% acceptance, roughly 2% reply, and 0% ICP approval. A blended average would have hidden both. Once I could see it, I did not write a note about it, I wrote a gate: the system now computes acceptance per source cohort and refuses to send for any cohort below a 30% floor once it has at least 15 resolved samples, with a brand-new-source grace period and an allowlist override.

What it cost: the losing lane was over half my sourced volume. Killing it meant a visibly smaller pipeline while the good lane refilled.

Source: per-lane funnel figures in outbound/auto-outbound/v1-system-map.md §4, measured. Gate implemented in executions/connection_cohort_gate.py (acceptance_rate_floor default 0.30, cohort_min_sample 15). My own phrasing of the belief, from job-hunt/gtm-vocabulary-crib-sheet.md: "If a lead source is producing ICP matches that never accept, I want that killed on measured numbers at n of 30, not on a hunch six weeks later."

03

Anything irreversible gets gated in code, never in a document

A stored approval must never quietly become permission to send later.

+

Written process is advice. A flag is a rule. In my system the scheduled DM job computes DRY_RUN as true unless an explicit --execute-approved-send argument is present, so the default state of a live send path is "do not send". Paid enrichment calls need a second, separate spend flag. Sourcing and cleanup scripts raise a hard exit rather than proceeding without their confirm flag. And every write to the lead table goes through one function that appends who wrote it, why, and what changed to an audit log, whether or not the API call succeeded. Approval discipline that lives only in a runbook survives exactly until the first busy week.

What it cost: speed. Every genuinely intended send needs an extra deliberate action, and I have argued with my own gates at 11pm.

Source: executions/cron_send_dm1s.py lines 36-37 (verified); single-exit writer executions/nocodb.py apply_lead_update(), audit trail at executions/logs/lead_writes.jsonl; hard exits in signal_web_sourcing.py, linkedin_activity_scrape.py, signal_web_import.py, cron_invite_cleanup.py.

04

Cheap models for volume, expensive judgment for decisions

136 research agents, roughly 1.1 million tokens, zero usable output.

+

I tried running one high-end research agent per lead. It burned around 1.1 million tokens across 136 agents and produced nothing I could use. I rebuilt it: research now runs as a standalone pipeline on a cheap model at roughly four to six cents per lead, with a deterministic schema validator deciding whether the output is acceptable, and the expensive model reduced to orchestration and the approval gate. The same discipline shows up in the signal-play generator, where enriching 78 leads cost 468 searches and 37.8 cents of model spend against a 50 cent cap. Unit economics per qualified output is a GTM number, not an infrastructure footnote.

What it cost: per-lead depth tuning. Every lead now gets one consistent research contract instead of a bespoke one, which is worse for the best leads and much better for the pipeline.

Source: incident and rebuild recorded in .claude/memory/feedback_research_architecture.md; enrichment run costs in outbound/cold calling/lead lists/live/2026-07-19_myphoner_210854_callable_enriched_summary.json.

05

Do not rewrite the script on fourteen calls

I had the pattern, I had the rebuttals drafted, and I parked them.

+

A review of 14 recorded calls found one objection appearing in 3 of them and five of the fourteen dying at the gatekeeper. I drafted the rebuttals and the line-level script edits, then deliberately did not apply them, with the reason written into the file: the sample was too thin to rewrite objection handling on. They sat behind a promotion gate that required the pattern to repeat across a fuller week before touching the live script. The discipline that makes a testing culture real is not running the test. It is refusing to act on a result you like.

What it cost: a month of not fixing something I had already diagnosed, and the rebuttal changes are still unpromoted.

Source: outbound/cold calling/call reviews/2026-06-30-cold-call-loser-analysis.md (14 recordings, 3/14 and 5/14 counts, measured); staging and promotion gate in pending-script-changes.md, header "DRAFTED, NOT APPLIED... Sample was only ~14 calls, too thin to rewrite objection handling on".

06

Personalisation has to be evidence, not intimacy

"Bro, this doesn't make any sense. Do you have something you actually want to say?"

+

That is a real reply to a message I sent. I had opened on the prospect's side project rather than the job he actually does, and referenced his site in a way that read as a scrape rather than a read. The same day, a message that opened with sector-specific jargon only someone who had genuinely looked at that supply chain would use got a positive reply. The lesson is not "personalise more". It is that a detail proves attention only when a stranger could not have produced it cheaply. Fake specificity is worse than no specificity, because it announces the automation. I also banned synthesised or recreated proof images outright: a screenshot is either a real capture of a real thing or it does not go.

What it cost: image coverage. Leads with no capturable signal now get a message with no image rather than a weaker one, which looks worse and converts better.

Source: both replies and the four-part root cause in .claude/memory/feedback_dm_failure_patterns.md; the no-synthesised-proof rule in .claude/memory/feedback_signal_screenshot_tactic.md; enforced at send time by QA rule R4_non_signal_image_path in executions/qa/rules/dm_to_reality_v1.md.

10 · Evidence, in progress

The same thinking inside the tools a GTM team already runs.

These four are live placeholders, being added over the coming weeks. I would rather show an honest empty frame than a stock screenshot of somebody else's table.

Slot 01 · Clay

Waterfall enrichment, built and annotated

Provider order, fallback logic, and the point at which I stop paying for a contact. The judgment I want visible is not how to build a waterfall, it is when to stop spending on one.

Screenshot + short annotation. Coming.

Slot 02 · Clay

Signal scoring and ICP fit in a table

The same scoring logic I run in Python, expressed in Clay: trigger detection, source capture on every scored row, and the threshold that decides whether a row is allowed to become a message.

Screenshot + scoring formula. Coming.

Slot 03 · Walkthrough

One play, end to end, on video

A single signal followed all the way through: trigger fires, entity resolves, row scores, message drafts, gate holds it, human approves, send, reply classified, outcome written back to the trigger that caused it.

Loom, roughly 6 minutes. Coming.

Slot 04 · Walkthrough

The gate refusing to send

The least glamorous and most useful demo I can give: the system declining to act because a flag is absent, a cohort is below floor, or a claim has no source. Safety is easy to assert and boring to prove, so I would rather prove it.

Loom, roughly 3 minutes. Coming.

On Clay specifically: I build in Clay, and the logic underneath it (enrichment waterfalls, signal-based targeting, segmentation, specialised routing) I built for myself in Python first, running daily on my own pipeline. That is why I come at it knowing which signals are worth chasing and which enrichment is worth paying for, rather than which buttons to press. Graded as claimable in my own audit file, job-hunt/references/claim-boundaries.md.

11 · Owned GTM

Rent the commodity parts of a GTM stack forever. The parts that are your actual advantage should end up as code you own.

This is the view most likely to make a company either want me or not want me. Both outcomes are useful, so here it is early rather than in month two.

The strongest argument against me, which I hold too

Ownership is stage-dependent, and before product-market fit, speed beats it outright. A team that builds its enrichment layer at $1m ARR has bought a maintenance burden instead of a moat. It is also worth noticing that nearly everyone publicly arguing "own your stack" sells owned infrastructure, which is a reason to discount the argument, including mine.

So I sequence it. Plug-and-play tools go in first so the motion can be proven cheaply, and only the parts that demonstrably carry the advantage get reverse-engineered in-house afterwards. Building before you have proof is how a GTM team accidentally becomes a platform team. Rented tooling is not the enemy. Renting the thing that makes you different is.

Source: the front-load-tools-then-reverse-engineer sequencing is stated in my post-interview debrief record as the explicit design of the 30/60/90 it accompanied. The stage-dependence counterargument is the mainstream position in current build-versus-buy writing and I have not seen it convincingly refuted.

I have already done it to myself, which is why I know when not to

I replaced parts of my own enrichment stack with my own scripts, wrote my own LinkedIn automation layer including the rate limiting and detection controls rather than renting a sending tool, and run my system of record on self-hosted infrastructure on a box I control. That is also why I can be useful about when not to. I know what each of those cost me to build and to keep alive.

Source: executions/unipile.py, 1,492 lines including a six-gate pre-flight check, Gaussian send delays and session breaks; self-hosted datastore and scheduling per outbound/auto-outbound/directives/crontab.md. Verdicts audited in job-hunt/references/claim-boundaries.md. The thesis as I sent it to a hiring principal after an interview is summarised in job-hunt/[redacted]-interview/post-interview-debrief.md lines 103-107, alongside the build-versus-buy reference I cited with it (amplemarket.com/blog/build-vs-buy-your-gtm-stack). [UNSOURCED — needs verification] the verbatim prose of that email was sent from a personal inbox and is not stored in the repo; the wording above is a faithful restatement, not a quotation.

12 · Contact

The interview offer stands.

Book the interview and I will bring the 30/60/90 and the pipeline diagram for your company, built before we speak. If you disagree with section 11, bring that too. It is a better first conversation than a competency question.