Training and reference implementation

Governance for Agent Engineering Teams

Your team adopted agentic coding tools. Now background agents open pull requests faster than anyone can review them, and nobody can say with certainty which agent did what, under whose authorization, or whether the checks it reports green are the checks that ran. Info Science AI teaches a governance framework that answers those questions by construction. It was built by a practitioner who runs it in production, and it comes with a reference implementation your team can adopt.

Book a session See the offerings

Who it is for

Teams already running agents

Engineering teams running Claude Code, Copilot coding agents, Cursor, Devin, or similar tooling in their CI/CD pipeline, with real work flowing through it, who have already seen an agent claim a green run it did not earn.

Who it is for

Leads who own the outcome

Engineering managers and technical founders who are accountable for what merges, know that review is now the bottleneck, and need an audit trail that would hold up if someone asked how a change got there.

Who it is for

Teams after an incident

An agent pushed something it should not have, or reported work it did not do. You want to know what actually happened and what to change so it cannot recur silently.

What you will learn

Governance enforced at the tool level, not in prompts

Instructions in a prompt are advice the agent may or may not follow. The framework taught here is deterministic and tool-agnostic: the controls live in the pipeline, not the prompt, they do not depend on which agent or which vendor you run, and they cannot be skipped. They produce a record rather than an assertion. Its defining property is that reward hacking is impossible by construction: an agent cannot grade its own work, cannot move the bar, cannot narrate its way past verification, and cannot edit the record. Seven elements. This is what they are; the workshop is how to build them.

  1. 1

    State outside of context

    The record of the work does not live in the agent's context window or memory. When a session ends or a context resets, nothing that matters is lost.

  2. 2

    Lanes and gates

    Work is isolated into governed streams with explicit gates. Where the work actually is can always be proven, and a crash mid-run does not lose it.

  3. 3

    Session identity and attribution

    Every write carries who made it and when, stamped by the tooling. Nothing is written anonymously, and two seats cannot silently overwrite each other.

  4. 4

    Fixed roles and stereo verification

    Agents may run tests. They may not grade them. No component grades its own work, and verification is arranged so that agreement is evidence rather than echo. Self-grading is impossible by construction.

  5. 5

    Sealed pre-registration

    The bar for a test is fixed and sealed before any data is read, by a party that never runs the test. An agent cannot grade on a curve it drew afterwards.

  6. 6

    The determination ledger

    Every determination is recorded in a space no single party controls, and the record can be checked on its own. History cannot be quietly rewritten.

  7. 7

    Gated hooks

    Code does not merge without explicit instruction, enforced by gated hooks in the pipeline rather than by a person watching. Standing prohibitions, each learned from a real incident, are built into the gates. Nobody babysits the agents; the structure does.

How each element is built, the files that make it up, and the mistakes that shaped them are the content of the workshops and the template pack below.

Flagship workshop

The outline

Eight parts. The half-day format covers the same ground in less depth; the recorded course follows the same order.

  • Why agent work needs governance, and why prompt instructions are not governance
  • State outside of context: manifest, lanes, rendered handoff
  • Lanes, gates, and crash recovery: the in-flight artefact
  • Session identity and attribution as chain of custody
  • Role separation: designing self-grading out of the system
  • Sealed pre-registration and the determination ledger
  • Gated hooks on commit and push, and standing prohibitions learned from incidents
  • Working session: map the attendees' current setup against the framework
Template pack

What you leave with

Included with the flagship workshop. Built from what runs in production, adapted to a generic layout so your team can adopt it as is or change it.

  • Manifest: the single source of truth for a project, with the fields that keep agents and humans aligned.
  • Lane schema: owner, state, gate log, in-flight artefact path.
  • Gate definitions: what must be true before a lane advances, and who may say so.
  • Pre-registration form: the sealed bar, with the fields that make a test decidable before it runs.
  • Ledger format: the determination record, and the check that verifies it on its own.
  • Reference tool: one dependency-free Python file that stamps, seals, and verifies.

The client implements. Info Science AI teaches and hands over the templates.

Offerings

Six ways to work with us

Priced for small and mid-size engineering teams.

Start here

Recorded course or short live session

$300 to $500 per seat

For individual engineers and small teams who want the framework explained end to end, on their own schedule.

  • The seven elements, walked through with the reference tool
  • Recorded, per seat, watch at your pace
  • Live format adds a question and answer session
Teaching
One team

Half-day team workshop

$1,200 to $1,500 per company

Live and remote, one company, unlimited developers. The outline compressed to a half day.

  • The eight-part outline, condensed
  • Your current setup mapped against the framework
  • Written summary of what was covered and what to look at first
Teaching
After you build

Implementation review

$1,500 to $2,500 per review

You set it up. We read what you built and say where it is thin.

  • Read of your manifest, lanes, gates, and ledger
  • Written findings: what is present, what is worth examining, what we would change
  • Follow-up call to walk through the findings
Advisory, not certification
Something went wrong

Incident review

$2,000 to $3,000 per incident

An agent pushed something bad, or claimed work it did not do. We reconstruct what happened from your records.

  • Timeline reconstruction from manifest, logs, and version control
  • Written account of what the record supports and what it does not
  • Recommendations on what to change so it cannot recur silently
Reconstruction, not assurance
Why Info Science AI

A running practice, not a curriculum assembled from reading

The framework taught here was developed against a codebase of roughly 250,000 lines maintained by successive generations of agent workers, then transferred upward to production AWS stacks running across many containers, where it works better, because more moving parts means more places for state to drift.

It governs engineering work for a financial engineering firm with six figures of live capital riding on the output. Every gate, prohibition, and ledger rule exists because something once went wrong without it.

One case: our precision audit caught a drift of up to 1.66 × 10−4 in a derived constant before it reached production, and we closed the class rather than patching the instance. The firm's own CTO, a PhD in mathematics, then ran all four verification modules independently and every assertion passed. The full account is on the Secure AI page.

Greg Hutchins, creator of the Certified Enterprise Risk Manager framework and author of more than 30 books on engineering, risk, and quality, said it on the record: "If people need consulting in AI assurance or AI governance, go to Daniel." The episode is on the AI Think Tank Podcast.

Underneath all of it is twenty years of evidence discipline. Before AI, Daniel Wilson ran an independent third-party testing lab and worked as a digital forensic specialist. The sealed bar, the ledger, and the chain of custody on every write come from that work.

Common questions

Before you book

Do you certify or sign off on our system?
No. Nothing we deliver is an audit, certification, or compliance opinion. We teach the framework, describe what we see, and tell you precisely what the framework cannot catch. A verification that does not state its limits is a claim, not a verification.
Do we need Claude Code?
The reference implementation runs on Claude Code. The framework itself is tool-agnostic: it is about state, roles, gates, and records, and it applies to Copilot coding agents, Cursor, Devin, or anything else that opens pull requests in your pipeline.
Is any of this remote?
All of it. Workshops and reviews are live over video. The course is recorded.
What do we need to prepare?
Access to whatever your agents currently write to: repositories, task trackers, logs. The working session maps that against the framework.
Book a session

Tell us how we can help

A few lines is enough. We reply within two business days with a proposed format, a date, and a price inside the ranges shown above.

Or email dan@infoscience.ai directly.

Everything on this page is training material and a reference implementation, not an audit, certification, or compliance opinion.

Our Partners