Lab

Governance layer for marketing agents

Every action a marketing agent takes (send an email, publish a post, raise a budget) lands in a ledger and gets checked. The gate then records whether the action would run, wait for a human or be stopped. Each week the system reports which agents drifted from their usual tone.

Jump to policy controls

Architecture

Agents act; the layer decides what runs, waits or stops

Where the governance layer sits Marketing agents send every action to the governance layer. It asks a human approver about risky actions, and the approver ships the approved ones to the ad and email platforms. Clean actions pass straight through to the platforms. Marketing agents Zapier · Lindy · bots send · post · spend Governance layer checks · gate · log drift · weekly report Human approver admin or legal approve or reject Ad & email platforms ads · email · CMS social networks sends asks ships passes clean actions
Where the governance layer sits Marketing agents send every action to the governance layer. It asks a human approver about risky actions, and the approver ships the approved ones to the ad and email platforms. Clean actions pass straight through to the platforms. Marketing agents Zapier · Lindy · bots send · post · spend Governance layer checks · gate · log drift · weekly report Human approver admin or legal approve or reject Ad & email platforms ads · email · CMS social networks sends asks ships passes clean actions

In the source build the gate runs in dry-run: agents still act directly, and the layer records what it would have held. A pre-flight endpoint runs the same checks before an action ships.

Inside the layer

Agent governance pipeline Connectors send actions to a normalizer, which writes to an append-only event ledger. The ledger feeds the checks, which score each action for the approval gate. The gate queues held actions for a human; human overrides are written back to the ledger as new rows. The ledger feeds tone scores to the drift detector. Outcomes, drift flags, overrides and held actions roll up into the weekly report and digest, which the dashboard shows. Connectors Zapier · Lindy HubSpot · fixture Normalize one row per action + who acted, how sure Event ledger append-only PGlite (Postgres) Checks tone · claims competitor · spend regulated · conflict Approval gate auto · review · block dry-run policy rules Human override admin or legal latest row wins Drift detector daily tone per agent vs. own baseline Outcomes sessions, revenue joined by subject Weekly report + digest totals · top 3 · findings · drift alarms gate queue · outcome impact Dashboard Next.js, role-gated sends writes feeds scores queues writes feeds tone joins flags logs holds shows
Agent governance pipeline Connectors send actions to a normalizer, which writes to an append-only event ledger. The ledger feeds the checks, which score each action for the approval gate. The gate queues held actions for a human; human overrides are added back to the ledger as new rows. The ledger feeds tone scores to the drift detector. The ledger's counts and the drift flags roll up into the weekly report and digest, which the dashboard shows. Connectors Zapier · Lindy · HubSpot Normalize one row per action, who acted Event ledger append-only PGlite Checks tone · claims competitor regulated spend · conflict Approval gate auto, review or block Human override admin or legal latest row wins Drift detector tone per agent vs. own baseline Weekly report + digest totals · top 3 · findings drift alarms · gate queue outcome impact by subject Dashboard Next.js, role-gated sends writes feeds scores queues adds feeds counts flags shows
Event ledger: the only store, append-onlyData flowHuman override, written back as a new row

Nothing in the ledger is edited. Checks, gate decisions and human overrides are new rows, so any week can be replayed and audited from the rows alone.

Ledger replay

22 of 27 actions run; 3 wait for review, 2 blocked

27 fixture actions from four agents over ten days, rechecked under the policy in the controls: change it and every decision is recomputed. Select a row to see its full check trace. The tone, claims, competitor and regulated-language checks are summarised here. The brand & claims gate covers copy checks in more depth on text you paste in.

–Events in ledger
–Findings

Event log, oldest first. Select an event for its check trace.

    Check trace

    Select an event above.

    Approval gate

    5 actions wait for a human

    Actions the policy marked Needs review or Blocked. When you approve or reject one, you append an override row to the ledger, and the system decision stays beside it. The counts, the report and the revenue figure below all update.

    Acting as

    In the source build only these two roles can override a decision.

    –Override rows
    –Revenue on held actions

    Held actions are those whose current decision is Needs review or Blocked. The revenue figure is their fixture outcome revenue.

      Override log

      No overrides yet.

      Drift

      Lifecycle Bot · Email: one day off its usual tone

      For each agent and channel: the daily mean tone score. A day raises an alarm if it falls below the warn line, or if it drops further below the stream's own baseline than the drop setting allows.

      Agent and channel

      Flagged days

      Agent / channelDayScoreBaselineReason

      Weekly report

      A weekly digest with no model in the loop

      Built from the current ledger state, in the same structure as the source build's report and digest.

      No language model writes this digest, here or in the source build. It is assembled from counts and rule rationales. The only model call in the pipeline is the tone-of-voice judge: Claude in production when switched on, otherwise the rule-based mock judge this page uses.

      How it works
      Method
      Rule-based checks, a policy-driven approval gate, per-stream tone drift
      Built from
      agent-governance, a fixture-only build I ran through spec-driven autonomous build loops. About me
      Data
      Fixture data for a fictional company, Northwind. Not client data.
      Runs
      Plain JavaScript in this page. No libraries, no server.

      The data

      27 agent actions from 1 to 10 August 2026: 16 emails from a Zapier "Lifecycle Bot", 9 blog posts, social posts and ad budget changes from a Lindy "Blog Agent", and one budget change each from a "Growth Agent" and a "Budget Bot" on the same campaign. Six problems are planted: a shouting email, an unbacked revenue claim, a post naming competitors, a "risk-free" promise, a budget over its ceiling, and the two-agent budget collision. Outcome data (sessions, conversions, revenue) covers 22 of the 27 actions. Competitor names and URLs are swapped for fictional ones; everything else matches the source fixtures.

      The checks

      • Tone of voice: a rule-based judge starts at 1.0 and subtracts for banned phrases (0.25 each), upper-case shouting, exclamation marks, long sentences and emoji. Below the fail line fails, below the warn line warns.
      • Claims: words like "guaranteed" or "double your" fail when they touch revenue, savings, performance or health and no customer or study is named.
      • Competitor: naming any competitor on the never-mention list fails.
      • Regulated: phrases like "risk-free" or "bank-level" warn and need Legal.
      • Spend: an ad budget above the agent's ceiling fails. So does an agent sending more emails in one UTC day than the limit.
      • Conflict: two different agents acting on the same campaign or subject inside the window each get a warning that names the other.

      The gate

      Each warn or fail takes the first approval rule that matches its check type and verdict. The strictest decision across the event's checks wins: block over review over auto. A human override is its own row. The latest override for an event counts, and the system decision is kept beside it.

      In the source build the gate runs in dry-run mode. Agent actions are read from the tools' logs after they happen, and the gate records what it would have held. A pre-flight endpoint runs the same checks and the same decision on a proposed action before it ships.

      Drift

      Tone scores are averaged per agent, channel and day. A stream's baseline is the mean of its daily means. The source saves it the first time drift is read, and a scheduled job rolls it forward. This page keeps nothing between loads, so every read computes the baseline fresh.

      The replay

      Each step re-runs every check over the events ingested so far, the way the source's batch run scores a seeded ledger. That is why the first half of the budget collision only gets its warning once the second agent acts. With all 27 events and the default policy, every check verdict, score, gate decision, drift alarm and digest line matches the source repo's own run.

      In production

      The source is TypeScript on Next.js 15 (App Router), with Drizzle ORM over PGlite, an embedded Postgres. The dashboard has routes for the timeline, findings, per-agent scorecards, approvals, objections, digest, trends and outcomes. Four demo roles control access, and only admin and legal can override a decision. The digest goes to a file by default, or to Slack. I built it in phases from four fixed specs, one commit per milestone. Nothing was committed until the repo's verify script passed: typecheck, unit tests and production build, with smoke and end-to-end tests added along the way.

      Limits

      • Fixture-only. The live Zapier connector and the Claude judge are wired but switched off, and no real account was connected.
      • The mock tone judge is a heuristic. It catches shouting and banned phrases, not subtle off-brand writing.
      • Agent identity is self-declared. Attribution confidence comes from whether the tool sent an agent key, not from a verified credential.
      • Ten days with one or two emails a day is a thin baseline. One bad email moves a daily mean.
      • The outcome data has a correlation planted on purpose. It proves the join works, not that flagged copy converts worse.
      • Not shown here: objections between agents, policy ingest from the brand guideline, and the trends view. All three exist in the source.