I find the moment a company becomes a buyer, in public data, then build the machine that acts on it.
01 · The interview offer
Two things, built from public sources before we speak: a tailored 30/60/90 for the actual role, and an Excalidraw map of your pipeline from signal to signed deal. Both are wrong in places on purpose, because the fastest way to learn a company is to be corrected on something specific.
The diagram, from a real one · built before a real interview
Open full size ↗
Built for a capital-advisory firm ahead of a GTM Engineer interview, from their published sector list and public site copy only. Every trigger on it maps to a named public dataset. Nothing on it is confidential.
Source: the diagram is generated from a checked-in Excalidraw scene and a written content spec, both in my repo. Redaction note: the company name has been removed from the diagram and from every repo path on this page, because that process may still be live. The thinking is the point; the client is not. Paths appear here as job-hunt/[redacted]-interview/….
02 · The first 90 days
Learn what a good deal looked like before it arrived, make what survives repeatable, then let the unit economics decide what continues.
Ends with: a funnel baseline, a back-test naming which triggers actually preceded good deals, a working entity spine, one approved cohort, one reliability fix.
Ends with: two or three signal families running as a monitored loop, provenance and suppression enforced in the pipeline, and a written record of which hypotheses were kept and which were killed.
Ends with: signal families ranked by cost per qualified conversation, a documented sourcing spine with runbooks, a live experiment backlog, and a plan that says what to stop as clearly as what to scale.
The move most signal projects get backwards. "Take past engagements and ask a single question of each: was there a public trigger visible before they arrived, and how early? This inverts the usual order. Most signal projects pick plausible triggers and hope; I would rather learn which triggers actually preceded deals [they] already value."
The unglamorous part that decides whether any of it works. "A filing entity name rarely matches the operating brand, which is where most signal work quietly fails. Filing to company to domain to founder to contact, each hop carrying source, date and confidence. That is the difference between a signal list and a usable pipeline."
What I refuse to promise. "Baselines are agreed after observing the current system, not invented in advance. Nothing in this plan promises a number I have no basis to promise."
Source: job-hunt/[redacted]-interview/30-60-90-content.md, lines 45, 47, 112, company specifics removed. A second, independently built roadmap for a different company exists at job-hunt/interviews/[redacted]/30-60-90-seed.md, which is how I know the method travels rather than being one lucky document.
03 · The numbers
Outbound is the system with the longest run behind it, so these are its numbers. Paid ads, content, AI search and client management have their own.
5% → 15% reply rate
Multichannel reactivation, around 200 leads a month, roughly 5 meetings a week booked. The client credits it with £3m+ in deal flow generated.
~20% response rate
Held across roughly 400 outbound touches a week, with 3 to 4 booked calls a week. The client now resells the system to their own customers.
10–20 calls a month
From around 100 leads a month, at a high close rate on a high-value legal service. Hours of weekly prospecting research removed from the team.
~4,000 emails a month
Roughly a 3% reply rate at handover, then operated in-house by the client without me. Built to be handed over, and still running after I stepped out.
29 clients closed
An outbound pipeline I sourced, booked, and closed personally, roughly £140k in lifetime value, alongside a 41% rise in new-client onboarding over six months.
Up to ~10% more firefighters
The client's own credited figure, from recruitment-ops automation that raised intake capacity, so more crews could deploy, including through the January 2025 California wildfire response.
Status: client-reported and client-credited figures from delivered engagements. Measured outcomes, not projections, and not the packaged system I sell today. Why they are here: an agency or founder background usually reads to a hiring manager as execution without infrastructure. This is the execution half. The next four sections are the infrastructure half, which is the part that is harder to fake. Full context on how each was measured is on the main portfolio, or ask me and I will walk you through any of them.
04 · GTM systems · the signal layer
The next four sections are one system, in four layers: the signal spec, the engine, the experiment loop, and the call loop. It runs on my own pipeline every weekday. Not a portfolio piece, which is why it has monitoring, audit logs, a post-mortem, and a gate that has told me no.
Everyone says "signal-based". Most of what gets called a signal is a description of a company, which is useless, because it was true last year and will be true next year. The rule above does most of the work. The second rule does the rest: every signal has to answer "why does this company need what I sell, now?"
| Trigger | Publicly observable in | Why it means "now" |
|---|---|---|
| Mid-raise, round not filled | SEC Form D, item 13.c | Item 13.c is literally "Total Remaining to be Sold". A dated public record of a company actively raising and short of target. |
| 16 months post-seed, no Series A | round-date data | Median seed to Series A gap is around 20 months. Past that point, the clock is the pressure. |
| First-ever CFO or Head of Finance posted | job boards, LinkedIn | The first finance hire is a company admitting money has become the constraint, usually before it says so anywhere else. |
| Contract award larger than anything delivered before | SAM.gov, USAspending, Find a Tender, TED | Delivery has to be funded, planned or not. Comparing an award against that same supplier's own prior awards needs no balance sheet, which is why it works across jurisdictions. |
| Clinical phase moves up | ClinicalTrials.gov | A phase transition is a dated, public commitment to a materially larger spend. |
Source: the signal set from the diagram in section 01, sources as drawn. Median seed-to-Series-A figure cited in my spec to Carta Q2 2025 (616 days). These are the triggers for one specific buyer. The transferable part is the shape: an event, a named dataset, and a distinct reason for urgency.
On rule 1. "Company is profitable" is a trait. "Company filed accounts showing a third profitable year" is a signal, because the filing happened on a date and is publicly observable. Traits give you a list. Only events give you a reason to call this week rather than any other week.
On rule 2. If four triggers all mean the same thing, that is one trigger with four sources. A good signal set is a set of genuinely different reasons the same problem becomes urgent. That is the difference I look for when someone shows me their targeting.
Source: both rules are written as governing constraints in my own diagram content spec, under "Governing rule for signals". They exist because an earlier version of that diagram was rejected for putting a trait on it, and I wrote the rule so it could not happen again.
A signal I cannot point at is a guess with better branding. In my own system a signal scored above 7 has to carry a source URL or the record fails validation and never reaches a message. It is a schema rule, checked by code before anything is written back.
Some signals point the wrong way. "I treat hiring as an anti-signal in my own lane. A company hiring a GTM Engineer is not a buyer, it is a competitor forming." Knowing which triggers to invert is the same skill as knowing which to chase, and it is cheaper to learn once than to discover across three months of sends.
Source: both stated verbatim in job-hunt/gtm-vocabulary-crib-sheet.md; the source-URL rule is enforced by outbound/auto-outbound/executions/research_schema_validate.py, which is deliberately code-only and does not judge content quality.
05 · GTM systems · the engine
73 Python scripts and 24,327 lines across 8 scheduled jobs, 304 commits since December 2025, and 3,156 leads sourced and evaluated all time. Three stages, each one gated.
Scrapes and imports from several lanes, then a two-pass ICP filter: deterministic pre-filters first (geography, freelancer patterns, competitor names, script-ratio checks) to strip obvious rejects before any model call, then a scored model pass. Every source is scored across the funnel separately, and a cohort gate blocks sending for any source under its acceptance floor.
icp_filter.py · source_scorecard.py · connection_cohort_gate.py
Each connected lead gets one deep-research pass on a cheap model, validated against a schema by code rather than by another model. The message layer generates variants, selects an image only if a real captured signal exists, and every draft passes a rule-based scanner with 15 blocking rules before it can reach an approval gate.
research_schema_validate.py · image_candidate_select.py · qa/rules/dm_to_reality_v1.md
Sends behind explicit flags with per-account daily caps and randomised volume. Replies are classified deterministically into eight classes, and only terminal classes can be auto-dispositioned. A state-legality rule pack blocks illegal transitions, such as advancing a follow-up sequence while a reply sits unhandled.
unipile.py · classify_reply.py · reply_routing.py · qa/rules/lead_state_consistency_v1.md
Source: file and line counts measured over outbound/auto-outbound/executions/**/*.py; lead count from outbound/auto-outbound/v1-system-map.md line 203, full-universe evaluation dated 2026-07-04; commit count from git rev-list --count HEAD against a first commit of 2025-12-20. Read the authorship disclosure below before reading those line counts as hand-typed.
Live sending stopped for three weeks and nothing told me. An acceptance-rate gate was computing over all time instead of a rolling window, and a partial deploy had shipped the scheduled jobs without the libraries they imported. The pipeline looked calm because a stalled system and an idle system produce the same silence.
What I built after: a deadman monitor that reads each job's own log tail, computes silence against a per-job tolerance, knows the difference between an expected quiet weekend and a real stall, and batches findings into one alert. Every scheduled job now writes a heartbeat even when it does nothing, so "no log line" can only ever mean the cron did not fire.
Source: incident and root cause in outbound/auto-outbound/v1-system-map.md §5b (2026-06-11 to 2026-07-01); monitor in executions/cron_healthcheck.py; heartbeat convention in directives/crontab.md.
The first version of my cohort acceptance gate would have reported close to 100% acceptance for every cohort, forever. Two independent 14-day windows interacted to create survivorship bias: connections that had not yet resolved were being excluded from the denominator in a way that only ever removed the failures.
I found it because the number looked too good, not because anything errored. The explanation, the fix, and a tunable resolution-lag constant are now written into the top of that file so the next person to read it cannot reintroduce it.
Source: executions/connection_cohort_gate.py module docstring and CONNECTION_RESOLUTION_LAG_DAYS.
Every one of these has real client code in the repo that makes real calls, rather than a logo on a slide: NocoDB as the system of record, Unipile for LinkedIn, Apify for profile and engagement scraping, Tavily and Exa for company-signal discovery and people resolution, Cal.com for bookings, OpenRouter and DeepSeek for qualification and research, Telegram for alerting, ReportLab for generated resource assets, and a self-hosted n8n instance alongside it.
I keep a separate audited file that grades my own tool claims as production, real-but-limited, or name-only, so nothing on a CV outruns what a reviewer would find in the code. I would rather hand you that file than have you find the gap yourself.
Source: integration list per file, in outbound/auto-outbound/executions/. The self-audit is job-hunt/references/claim-boundaries.md, produced by three independent agents reading the code rather than the CV.
I use Claude Code heavily and my commit history says so plainly. What I am claiming is the architecture, the failure modes, the safety model, and the debugging: the data model, the providers, the gating rules, the schema contracts, the approval policy, the failure handling, and the writeback path.
When the scraper died silently for three weeks, no model found that. When the acceptance gate started reporting numbers that were too good, no model found that either. I found both because I was watching the numbers and knew what they were supposed to look like. That is the job. The typing is not the job.
Source: pre-committed position recorded in job-hunt/references/claim-boundaries.md, "The AI-authorship boundary". The public reference repositories under my GitHub are demonstrations on fictional data and say so in their own READMEs; they are not client work and not this production system.
06 · GTM systems · the experiment layer
Any prompt surface in the system can register into a shared framework with a real stopping rule. Four constants govern the whole thing.
| Constant | Value | What it governs |
|---|---|---|
| Promotion threshold | 0.85 | Probability of superiority required before a challenger replaces the baseline, or the baseline kills the challenger |
| Monte Carlo samples | 10,000 | Draws per arm, per read, off each arm's Beta posterior |
| Volume floor | 10 per side | Below it the verdict is "not enough data" rather than a number |
| Registered surfaces | 7 (1 live) | Declared in YAML with table, reward field and thresholds. A new test needs no new code. |
Source: executions/optimization/significance.py, constants DEFAULT_THRESHOLD_PROB = 0.85, DEFAULT_MIN_PER_SIDE = 10, DEFAULT_SAMPLES = 10_000, all verified in file. Surface list and statuses from optimization/surfaces.yaml.
The maths itself is numpy and about two hundred lines, with a self-test battery of hand-checked cases in the same file. It replaced a fixed-sample heuristic that implied five to six weeks per decision, which at my volume made the loop useless.
optimization/route_variant.py · optimization/registry.py · optimization/surfaces.yaml · optimization/prompts/generator.md · qa/rules/dm_to_reality_v1.md
The current cycle tests message length, cutting the first message from 40 to 60 words down to 25 to 35 with a hard cap at 40. The baseline was recorded first: 45 sends, of which 36 were mature enough to count, producing 6 positive replies, a 16.67% positive-reply rate. The expected lift was declared before any data came back, at a floor of three percentage points and a realistic case of five.
Those lift figures are a stated hypothesis. They are not a result, and I will not present them as one until the posterior says so.
Source: outbound/auto-outbound/optimization/dm1/learnings.md. Baseline figures measured; expected-lift figures explicitly labelled in the file as expected floor and realistic, not outcomes.
07 · GTM systems · the call-data loop
Objection data is the highest-quality targeting input available, and it only exists downstream of the send. This is the reason I want the whole funnel rather than the top of it.
Recordings transcribed on my own machine with a local speech model. No per-minute billing, no audio leaving the box, no reason to ration how much gets analysed.
executions/transcribe_recordings.py (faster-whisper, CPU, offline)
Transcripts matched to CRM rows by content, dispositions pulled from the dialer's API rather than retyped. The review document says out loud that transcripts carry no speaker labels and roles are inferred, and records the one attribution error I made and corrected.
call reviews/2026-06-30-cold-call-loser-analysis.md, lines 4-6 and 58-60
Findings become drafted rebuttals in a separate staging file with the evidence quoted next to each one, gated on the pattern repeating across a fuller sample before it can touch the live script.
call reviews/pending-script-changes.md
The single most common objection across 14 recorded calls, appearing in 3 of them, was a version of "we already do this ourselves". Read as a copy problem, that is a rebuttal to write. Read as a targeting problem, it says something harder: the segment I was calling was capable enough to believe it already owned the thing I was selling.
A month later I cited that exact finding, by file and by count, as the reason a new script dropped that segment from its target market entirely and moved to a different buyer with a different, smaller entry offer. Call data changing the script is table stakes. Call data changing who you call is the version that moves revenue.
Source: objection count from outbound/cold calling/call reviews/2026-06-30-cold-call-loser-analysis.md (3 of 14, measured); the targeting change and its citation of that evidence in outbound/cold calling/playbooks/cold_call_script_v4_signal_feed.md, lines 28-29.
08 · GTM playbooks
Same test every time: an event with a date, a named public source, and a distinct reason it means now. Two of these I have run. Four are designed and ready. The mechanics are the claim, and I will defend any of them.
Status: these are plays, not case studies. "Run" versus "Designed" follows how each is written in my own notes: two are described in past tense as work I did, four in conditional as designs. Deliberately absent: the outcome figures that used to sit on these. They were projections, and a projection printed in past tense reads as a delivered result, so they are gone rather than relabelled. The mechanic is the claim. [UNSOURCED — needs verification] per-play outcome numbers.
09 · Principles
A principle with no cost attached is a slogan. Open any of them for the decision, the tradeoff, and the file.
10 · Evidence, in progress
These four are live placeholders, being added over the coming weeks. I would rather show an honest empty frame than a stock screenshot of somebody else's table.
Provider order, fallback logic, and the point at which I stop paying for a contact. The judgment I want visible is not how to build a waterfall, it is when to stop spending on one.
Screenshot + short annotation. Coming.
The same scoring logic I run in Python, expressed in Clay: trigger detection, source capture on every scored row, and the threshold that decides whether a row is allowed to become a message.
Screenshot + scoring formula. Coming.
A single signal followed all the way through: trigger fires, entity resolves, row scores, message drafts, gate holds it, human approves, send, reply classified, outcome written back to the trigger that caused it.
Loom, roughly 6 minutes. Coming.
The least glamorous and most useful demo I can give: the system declining to act because a flag is absent, a cohort is below floor, or a claim has no source. Safety is easy to assert and boring to prove, so I would rather prove it.
Loom, roughly 3 minutes. Coming.
On Clay specifically: I build in Clay, and the logic underneath it (enrichment waterfalls, signal-based targeting, segmentation, specialised routing) I built for myself in Python first, running daily on my own pipeline. That is why I come at it knowing which signals are worth chasing and which enrichment is worth paying for, rather than which buttons to press. Graded as claimable in my own audit file, job-hunt/references/claim-boundaries.md.
11 · Owned GTM
This is the view most likely to make a company either want me or not want me. Both outcomes are useful, so here it is early rather than in month two.
Ownership is stage-dependent, and before product-market fit, speed beats it outright. A team that builds its enrichment layer at $1m ARR has bought a maintenance burden instead of a moat. It is also worth noticing that nearly everyone publicly arguing "own your stack" sells owned infrastructure, which is a reason to discount the argument, including mine.
So I sequence it. Plug-and-play tools go in first so the motion can be proven cheaply, and only the parts that demonstrably carry the advantage get reverse-engineered in-house afterwards. Building before you have proof is how a GTM team accidentally becomes a platform team. Rented tooling is not the enemy. Renting the thing that makes you different is.
Source: the front-load-tools-then-reverse-engineer sequencing is stated in my post-interview debrief record as the explicit design of the 30/60/90 it accompanied. The stage-dependence counterargument is the mainstream position in current build-versus-buy writing and I have not seen it convincingly refuted.
I replaced parts of my own enrichment stack with my own scripts, wrote my own LinkedIn automation layer including the rate limiting and detection controls rather than renting a sending tool, and run my system of record on self-hosted infrastructure on a box I control. That is also why I can be useful about when not to. I know what each of those cost me to build and to keep alive.
Source: executions/unipile.py, 1,492 lines including a six-gate pre-flight check, Gaussian send delays and session breaks; self-hosted datastore and scheduling per outbound/auto-outbound/directives/crontab.md. Verdicts audited in job-hunt/references/claim-boundaries.md. The thesis as I sent it to a hiring principal after an interview is summarised in job-hunt/[redacted]-interview/post-interview-debrief.md lines 103-107, alongside the build-versus-buy reference I cited with it (amplemarket.com/blog/build-vs-buy-your-gtm-stack). [UNSOURCED — needs verification] the verbatim prose of that email was sent from a personal inbox and is not stored in the repo; the wording above is a faithful restatement, not a quotation.
12 · Contact
Book the interview and I will bring the 30/60/90 and the pipeline diagram for your company, built before we speak. If you disagree with section 11, bring that too. It is a better first conversation than a competency question.