I build GTM systems that run on AI.

I'm a marketing AI engineer. I build the measurement models, data pipelines and agents that marketing teams run on. Seven of them run live in the Lab.

Status

Available
immediately

  • Marketing AI engineering
  • Growth engineering & data

Selected work

A proactive GTM engine, built end-to-end

Scope
Signal pipeline → enrichment → scoring → branched outreach → eval loop
Stack
Python, dbt, Clay, LLM copy layer, CRM automation, GitHub/PyPI/Docker APIs

The challenge

Pipeline was 100% inbound: auto-scored, one manual reply each. It worked, but it was the whole motion. The brief was to build the next one.

What I built

An outbound engine that treats public developer activity as intent. It pulls every signal (stars, forks, installs, container pulls, telemetry, hiring), enriches each contact, and scores it on ICP fit plus signal strength.

  • Deduped daily against inbound, so no one gets double-touched.
  • A Python scoring module (stdlib only) makes every ranking auditable.
  • An LLM copy layer drafts each sequence, gated by deterministic rules before send.
  • Every conversion re-weights the model and grades the copy that worked.

The outcome

Shipped as a live app plus a runnable scoring model. Every number came from a real public API. Built to run on ~0.1 FTE.

6Signal sources unified
~14×Modelled year-one ROI
<5%Sends needing human review

Production multi-agent marketing system

Scope
Agentic infrastructure for a 30+ person marketing team
Stack
Claude / Claude Code, MCP (Canva, Linear, Miro), RAG, vector search, evals

The challenge

A 30+ person marketing team with zero agentic infrastructure. Everything ran on manual triggers.

What I built

A 24/7 system where orchestrator agents coordinate subagents in parallel. Content, data refresh, reporting, and knowledge upkeep all run without manual triggers.

  • Weekly agents propose rule changes; monthly cycles auto-approve them only if benchmarks improve.
  • RAG memory over a ~1,000-doc vector-indexed wiki, with semantic search and automated pruning.
  • Decay-weighted scoring that keeps high-value knowledge within token limits.
  • Eval frameworks: skill-trigger accuracy, output quality, and retrieval precision, with regression tracking.

The outcome

24/7Autonomous operation, no triggers
~1,000Docs in vector-indexed memory
30+Person team served

Decentralized multi-agent swarm, no orchestrator

Scope
Holacracy-inspired framework, no central orchestrator, reusable across use cases
Stack
LangGraph (StateGraph, Send fan-out), Claude Code CLI (headless), JSON Schema, Python

The challenge

The multi-agent system above runs on a supervisor: orchestrator agents decide who acts next. I wanted to test the other pattern: a decentralized model where no agent or human picks the next move. Would it hold together in production?

What I built

Independent roles coordinate through file-based mailboxes and a Holacracy-style governance mechanism: any role can file a proposal or an objection, and a proposal is accepted once every other role has had a turn to object and none did. Timestamps in an append-only audit log decide that, not an LLM or a person. LangGraph handles only the mechanical parts (state, parallel fan-out, routing). Every judgment call happens inside a headless Claude Code subprocess, and its output is schema-validated.

  • Found and fixed real permission-scoping bugs by testing directly against the CLI. One flag was silently overridden by a default; another blocked tools it claimed to support.
  • Built a local observability dashboard: cost-per-role tracking, an objection matrix, proposal-lifecycle metrics, all computed from structured logs at zero added inference cost.
  • The same framework runs multiple use cases via swappable per-pack role configs. A new use case needs a new config directory and no framework code.

The outcome

It runs on a live case: sourcing and screening real job postings and drafting tailored applications. A dedicated fact-checking role enforces a hard no-fabrication rule.

Step through a real run, live →

21Proposals in the replayed run, each decided by one mechanical rule
5Independent roles per pack, coordinating async
3CLI permission bugs found by testing, then fixed

Incrementality testing framework

Scope
Causal measurement framework, built from scratch
Stack
Python, ads-platform APIs, SQL, geo-experiments

The challenge

Spend ran on platform-reported conversions, numbers the platforms grade themselves on. No honest read on what was incremental.

What I built

A from-scratch incrementality framework in Python on the ads-platform APIs. It runs lift experiments that measure what spend caused.

  • Wired results into budgeting, so spend followed proven lift.

The outcome

~15%CAC improvement

GTM data infrastructure & conversion tracking

Scope
CDP ownership, server-side events, lifecycle tracking
Stack
Segment CDP, CAPI, GTM, lifecycle-stage events, GDPR-compliant pipelines

The challenge

Marketing, product, data, and legal ran on inconsistent signals. In-platform and backend numbers disagreed.

What I built

End-to-end Segment CDP, server-side CAPI events, and lifecycle-stage conversion events, all GDPR-compliant.

  • Closed the gap between in-platform and backend data with automated reporting.
  • Coordinated marketing, product, data, and legal to keep pipelines clean.

The outcome

2M€+ARR uplift contribution
~7%Conversion-rate improvement
~10%More open pipeline generated

Open source

Lab

Run the work yourself.

Seven tools from my work on AI agents and marketing measurement. Each one runs in your browser on sample data, and three also take your own CSV. Change an input and the results update.

All lab projects →