<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>Field Notes &#8212; Allyn L. Shaw</title>
  <link>https://allynshaw.com/notes/</link>
  <atom:link href="https://allynshaw.com/feed.xml" rel="self" type="application/rss+xml"/>
  <description>Working notes on AI governance, technology risk, and execution.</description>
  <language>en-us</language>
  <lastBuildDate>Mon, 03 Aug 2026 12:00:00 +0000</lastBuildDate>
  <item>
    <title>Decision Integrity</title>
    <link>https://allynshaw.com/notes/decision-integrity.html</link>
    <guid isPermaLink="true">https://allynshaw.com/notes/decision-integrity.html</guid>
    <pubDate>Mon, 03 Aug 2026 12:00:00 +0000</pubDate>
    <description>Decision quality is about the single call. Decision integrity is about what that call does to every decision after it.</description>
    <content:encoded><![CDATA[<div class="sec">
  <div class="k">01 &middot; Where this started</div>
  <h2>Every decision was defensible. The totality was a failure.</h2>
<p class="lead">Years ago, at a prior institution, I watched a capable senior leader make a series of decisions that were each, taken alone, reasonable. Sound analysis. Right stakeholders in the room. Defensible on the day. And when those decisions merged in the real world, they produced an outcome nobody had chosen and everybody had built.</p>
<p>The hard part came after. I had to walk back through the chain and explain how each decision had quietly narrowed the options for the next one — how a sourcing call constrained an architecture call, how the architecture call locked in an operating assumption, how the assumption aged badly and took three downstream commitments with it. No single decision was wrong. The system of decisions was.</p>
<p>That experience has stayed with me for a simple reason: nothing in how we run large organizations would have caught it. We review decisions one at a time. We reward decisiveness. We document outcomes. We almost never examine what a decision does to the decision space around it. This note is my attempt to name that gap.</p>
</div>
<div class="sec">
  <div class="k">02 &middot; The law</div>
  <h2>Decisions compound.</h2>
<p>Every consequential decision changes the conditions under which future decisions get made. It forecloses options, creates dependencies, sets precedents, and allocates attention. That's not a side effect — it's the primary long-run consequence, and it usually dwarfs the direct one.</p>
<p>Which means organizations accumulate <strong>decision debt</strong> the way codebases accumulate technical debt. Shortcuts in how decisions get made — unclear ownership, unstated assumptions, horizons that stop at the current quarter — don't fail immediately. They compound quietly, and the interest comes due later, paid by people who weren't in the room.</p>
<div class="pull">No single decision was wrong. The system of decisions was.</div>
<p>Here's the distinction the whole argument turns on: <strong>decision quality is not decision integrity.</strong> Quality asks whether this call was sound — right information, right analysis, right judgment. Integrity asks whether this call holds together with everything upstream and downstream of it: whether it honors the constraints it inherited and whether it's honest about the constraints it creates. An organization can be full of high-quality decisions and still be drowning in decision debt. That was exactly the situation I watched unfold.</p>
</div>
<div class="sec">
  <div class="k">03 &middot; The gap in the discipline</div>
  <h2>Decision Risk deserves its own layer.</h2>
<p>We've professionalized nearly every risk that matters to a large institution — credit, market, operational, technology, model, third-party. Each has owners, appetite statements, and controls. But the risk that a series of individually sound decisions compounds into an unsound position has no name, no owner, and no framework. It shows up in the post-mortem, never in the register.</p>
<p>I've started calling it Decision Risk, and I think it's the most under-managed risk in modern organizations — and getting worse. AI is collapsing the cost of analysis, which means decision velocity is going up everywhere. More decisions, made faster, by more actors, with fewer natural pauses. Compounding accelerates with volume. The organizations that thrive won't just make better individual calls; they'll manage the integrity of the whole decision system.</p>
</div>
<div class="sec">
  <div class="k">04 &middot; The five tests</div>
  <h2>What a decision should pass before it ships.</h2>
<p>The practical core of Decision Integrity is five tests. None require a committee. All require honesty.</p>
<div class="fm"><div class="tn">01</div><strong>The System test.</strong> Do we understand what this decision touches? Not the org chart — the actual web of processes, commitments, and prior decisions it lands in. A decision made against an imagined system fails in the real one.</div>
<div class="fm"><div class="tn">02</div><strong>The Ownership test.</strong> Does one person own this decision — and its downstream consequences? Not consulted, not aligned. Accountable. Decisions with distributed ownership have no ownership, and orphaned consequences become someone else's crisis.</div>
<div class="fm"><div class="tn">03</div><strong>The Time test.</strong> How does this decision age? Some calls are right for eighteen months and wrong at year three by design. That can be fine — if the expiry is stated. A decision with an unexamined shelf life is a liability with no maturity date.</div>
<div class="fm"><div class="tn">04</div><strong>The Dependency test.</strong> What does this decision force, foreclose, or assume? Every option it kills and every future call it pre-commits is part of its true cost. If you can't name the dependencies, you haven't priced the decision.</div>
<div class="fm"><div class="tn">05</div><strong>The Integrity test.</strong> Would this decision survive being seen whole? Alongside the ones before it, by the people who'll inherit it, with its assumptions written down. If it only looks good in isolation, it isn't good.</div>
</div>
<div class="sec">
  <div class="k">05 &middot; Deliberate debt</div>
  <h2>Taking on decision debt isn't the sin. Hiding it is.</h2>
<p>None of this argues for slow decisions or decision-by-committee. Speed matters, and sometimes the right call is to knowingly take on decision debt — lock in a vendor before the strategy is settled, ship the interim architecture, accept the precedent. Engineering teams do this with technical debt all the time, and the good ones do it well because they do it in the open: the shortcut is named, logged, and scheduled for repayment.</p>
<p>Decision debt deserves the same treatment. The failure mode isn't the shortcut. It's the unrecorded shortcut — the constraint nobody wrote down, discovered two years later by a team that can't understand why their options are so narrow. Deliberate debt is a tool. Silent debt is a trap.</p>
</div>
<div class="sec">
  <div class="k">06 &middot; The claim</div>
  <h2>Leaders are judged on the decision systems they leave behind.</h2>
<p>We evaluate leaders on results, and results matter. But results are partly luck and often lag. The more durable measure is what a leader does to the organization's capacity to decide well after they're gone: whether ownership is clear, whether assumptions get written down, whether debt is deliberate, whether the decisions they made left the option space wider or narrower for the people who follow.</p>
<p>That's the discipline I'm calling Decision Integrity. This note plants the flag; the full framework is in the works — how to score it, how to review for it, and what a decision-mature organization looks like in practice. If this names something you've lived, I'd like to hear the story.</p>
</div>]]></content:encoded>
  </item>
  <item>
    <title>Risk at Runtime</title>
    <link>https://allynshaw.com/notes/risk-at-runtime.html</link>
    <guid isPermaLink="true">https://allynshaw.com/notes/risk-at-runtime.html</guid>
    <pubDate>Tue, 28 Jul 2026 12:00:00 +0000</pubDate>
    <description>When software starts acting on its own, risk management built for quarterly review cycles stops working. The controls have to run at the speed of the systems they govern.</description>
    <content:encoded><![CDATA[<div class="sec">
  <div class="k">01 &middot; The lag is the risk</div>
  <h2>We measure last quarter. The systems act this second.</h2>
<p class="lead">I've spent three decades running the business of technology inside large banks, and for most of that time our risk tooling and our technology moved at roughly the same speed. Change was planned, released, and reviewed on cycles a human calendar could hold. A key risk indicator refreshed monthly was a reasonable proxy for reality.</p>
<p>That assumption just broke. Agentic systems — software that plans, decides, and acts — change their behavior in production, continuously. Against that, the standard apparatus of technology risk management is a set of lagging indicators: KRIs assembled after the fact, control tests run on samples, attestations signed quarterly about a world that no longer exists by the time the signature dries.</p>
<p>The uncomfortable truth is that the artifacts aren't the risk management. They're the exhaust of risk management as it was practiced when humans were the only actors. Keep producing them on a quarterly cadence while your systems act hourly, and the gap between what you report and what is actually happening becomes the largest unmanaged risk on your books.</p>
</div>
<div class="sec">
  <div class="k">02 &middot; The idea</div>
  <h2>Evidence computed, not assembled.</h2>
<p>The alternative is to move the point of control from the review meeting to the runtime. Concretely, that means three things.</p>
<p><strong>Controls become versioned, testable code.</strong> A control that lives in a policy document is an intention. A control expressed as policy-as-code executes on every relevant action, cannot be skipped by busy humans, and carries a version history the second line can challenge line by line.</p>
<p><strong>Every execution self-evidences.</strong> When the control runs in the path of the work, the evidence of its operation is generated at the moment of execution — signed, immutable, mapped to the obligation it satisfies. Nobody assembles an evidence package before an exam. The evidence already exists.</p>
<p><strong>Exposure is computed continuously.</strong> Residual risk stops being a rating negotiated in a workshop twice a year and becomes a number derived from live control performance, with drift visible the day it starts, not the quarter it's discovered.</p>
<div class="pull">No quarter-end scramble. The evidence already exists.</div>
<p>None of this weakens governance. It's the same obligations with stronger answers — and it's the only architecture I can see that lets a regulated institution adopt autonomous systems honestly, because it's the only one where the control layer runs as fast as the thing being controlled.</p>
</div>
<div class="sec">
  <div class="k">03 &middot; The transition</div>
  <h2>The dual rail: run both, retire one on evidence.</h2>
<p>You don't get to flip a switch on this, and you shouldn't ask anyone — least of all your second line or your examiners — to take a new system on faith. The transition I'd run is a dual rail over eighteen to twenty-four months: the traditional KRI and attestation rail keeps operating exactly as it does today, while the runtime rail comes up beside it, domain by domain.</p>
<p>The quiet weapon is the reconciliation log. For every domain on both rails, you record where the runtime evidence and the traditional artifacts agree and where they diverge, and you disposition every divergence. When reconciliation holds for consecutive cycles, you've built the evidentiary case to retire the manual rail for that domain — on data, not on trust. The transition has an owner and an end state. A permanent dual system is a failure, not a compromise.</p>
<p>Where to start matters. I'd begin on the agentic estate itself: it's where the gap between change velocity and assessment cadence is widest, and no legacy KRI regime has earned incumbency there. Prove the spine on the sharpest edge — a live inventory of agents and their permissions, autonomy tiers that are earned through evidence gates rather than granted by calendar, and one existing committee metric rendered directly from runtime data with full lineage. Ninety days to first value, then expand deliberately.</p>
</div>
<div class="sec">
  <div class="k">04 &middot; The regulatory translation</div>
  <h2>Same obligations. Stronger answers.</h2>
<p>For institutions under heightened supervisory standards, this is not a workaround — it's strengthened compliance. The examiner conversation changes from "show me your testing" to "here is the decision log."</p>
<div class="occ"><table><thead><tr><th>Supervisory expectation</th><th>Today's answer</th><th>Runtime answer</th></tr></thead><tbody><tr><td>Front line owns and demonstrates control effectiveness</td><td>Quarterly self-attestation, sampled testing, issues raised after the fact</td><td>Every execution self-evidences; the front line's controls prove themselves in production</td></tr><tr><td>Independent risk management challenges effectively</td><td>Second line reviews stale samples and challenges the testing approach</td><td>Second line challenges the policy itself — versioned, testable code — and monitors the full population</td></tr><tr><td>Timely escalation of breaches to risk appetite</td><td>Breach discovered at the next testing cycle; escalation measured in weeks</td><td>Breach blocked or flagged at execution; escalation measured in minutes</td></tr><tr><td>Talent and capacity commensurate with risk profile</td><td>Thousands of hours consumed producing and checking manual attestations</td><td>Those hours move to policy design, agent oversight, and hard risk judgment</td></tr><tr><td>Governance keeps pace with new activities</td><td>Agent activity governed by controls designed for human-speed change</td><td>Agents can't act outside policy — the control layer runs at the same speed as the agents it governs</td></tr></tbody></table></div>
</div>
<div class="sec">
  <div class="k">05 &middot; Honest constraints</div>
  <h2>Where this goes wrong — named in advance.</h2>
<p>I have more credibility proposing this if I name its failure modes myself. There are four I watch.</p>
<div class="fm"><div class="tn">01</div><strong>Quantification theater.</strong> Computed exposure numbers are only as good as the loss scenarios and frequency estimates beneath them. Start with a small set of scenarios backed by real event data, expand deliberately, and present precision honestly — ranges and trajectories, not false decimal points.</div>
<div class="fm"><div class="tn">02</div><strong>Automation overreach.</strong> An automated response that fires wrongly against a production payments flow is its own operational risk event. Response playbooks start conservative — alert and require approval — and earn their way to autonomous execution under the same discipline imposed on the agents themselves. The appetite framework applies to the risk system too.</div>
<div class="fm"><div class="tn">03</div><strong>Two systems of record.</strong> If the dual rail becomes permanent, you've doubled your cost and split your truth. The reconciliation log exists to retire the manual rail on evidence, and the transition period has an owner and an end state — not an indefinite extension.</div>
<div class="fm"><div class="tn">04</div><strong>The attestation shadow.</strong> The risks that resist telemetry — conduct, judgment, third-party concentration — must not become second-class citizens because they lack a live number. A verified-and-attested ledger keeps them visible, and committee attention is allocated by materiality, not by measurability.</div>
</div>
<div class="sec">
  <div class="k">06 &middot; The ask</div>
  <h2>Start where the gap is sharpest.</h2>
<p>If you run technology or technology risk inside a regulated institution, the question in front of you isn't whether autonomous systems arrive — they're already in your pipelines. The question is whether your control environment governs them or merely documents them after the fact.</p>
<p>My argument: pick the domain where the gap between change velocity and assessment cadence hurts most. Codify your top controls as policy. Instrument the evidence at execution. Run the dual rail, keep the reconciliation log, and let the data retire the old system domain by domain. Ninety days to the first control that proves itself in production. Everything after that is expansion.</p>
<p>Risk management at the speed of the business used to be an aspiration. For autonomous systems, it's the entry fee.</p>
</div>]]></content:encoded>
  </item>
</channel>
</rss>
