<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Formation Governance — Christa Burger</title><link>https://christaburger.com/tags/formation-governance/</link><description>Christa Burger is a CISO and VP of Cybersecurity — twenty-plus years across cybersecurity, risk, resilience, and governance in finance and technology. Building operating systems that build and reinforce trust.</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><managingEditor>christa@christaburger.com (Christa Burger)</managingEditor><webMaster>christa@christaburger.com (Christa Burger)</webMaster><copyright>© 2026 Christa Burger. All rights reserved.</copyright><lastBuildDate>Tue, 08 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://christaburger.com/tags/formation-governance/index.xml" rel="self" type="application/rss+xml"/><item><title>The First Time They Disagree</title><link>https://christaburger.com/blog/the-first-time-they-disagree/</link><pubDate>Tue, 08 Sep 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/the-first-time-they-disagree/</guid><category>Series: Formation Governance</category><description>What an embedded evaluator needs when the welcome meeting is over.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/plexus-network.jpg" alt="Abstract digital network of interconnected nodes and concentric rings representing oversight and connection" /></figure><p>Oversight is easiest to support before it changes a decision.</p>
<p>At the beginning of an evaluation arrangement, everyone can agree that transparency matters. The charter is being drafted. The right people have joined the meeting. Someone may even be getting a badge. The harder test arrives later, when the evaluator finds something material, the release matters, the team is tired, and the difference between &ldquo;we should understand this better&rdquo; and &ldquo;we can proceed&rdquo; has a price attached to it.</p>
<p>That is the meeting an oversight model has to survive.</p>
<p>In his September essay, Dario Amodei argues for embedding independent evaluators inside frontier AI companies, with continuing access to systems and people and rights to publish findings. He draws on a banking-supervision analogy, which is useful because supervision is not just a report. It is an operating relationship with access, escalation, independence, and consequences. <a href="https://www.darioamodei.com/essay/machines-of-loving-grace">We Must Pace the Frontier</a></p>
<p>The practical questions are not glamorous, but they decide whether the arrangement works. Access to what? In what format? With which logs, interviews, system records, and authority to ask follow-up questions? How does the evaluator know whether the evidence set is complete enough for the claim being made? What happens when the team believes the evaluator has misunderstood a test result? Who owns the decision if the disagreement remains unresolved?</p>
<p>There are ordinary ways to starve an evaluator who technically has access. The relevant data can be available but unusable. Logs can omit the behavior that matters. Every question can require an introduction to someone who is busy, traveling, or not quite sure who owns the current version. A badge opens doors. It does remarkably little about naming conventions.</p>
<p>METR has already named several important conditions for serious investigation: model and transcript access, employee interviews, adequate resources, communication with oversight bodies, and disclosure of how redactions affect conclusions. That gives the field something concrete to build on. <a href="https://metr.org">METR&rsquo;s investigation framework</a></p>
<p>The next step is to make those conditions visible with the finding itself. A useful report should show what investigators requested, what they received, what remained unavailable, and what each gap prevented them from determining. &ldquo;No evidence of a problem&rdquo; means one thing after direct access to the relevant behavior and another thing when the behavior was never logged. Sensitive evidence may need restricted handling; the public account should still explain how the restriction affected the conclusion.</p>
<p>The disagreement path matters just as much. If an evaluator identifies a concerning behavior and the lab believes it is an artifact of the test, that may be true. The right response is a documented competing explanation and a way to distinguish between them. Preserve the original observation, the alternative account, the evidence that would resolve the question, and the decision taken while uncertainty remains. A finding should be able to change because the evidence changed. It should also be possible to tell when the language changed because the meeting got uncomfortable.</p>
<p>Consequences need the same clarity. An embedded evaluator does not need unilateral control over company operations to matter. A serious finding does need a route to someone who can accept the risk, require a remedy, restrict an action, or explain why proceeding is justified. That decision needs an owner and a record. Otherwise the organization can comply with evaluation indefinitely while leaving the underlying condition untouched. The report becomes another artifact the system knows how to produce.</p>
<p>The most useful preflight exercise is simple: run a bounded disagreement before the stakes are existential. Give the evaluator an incomplete evidence set, a disputed finding, and a real escalation path. Watch how long it takes to reach the right people. Watch whether uncertainty survives the executive summary. Watch who can request more work, who pays for it, and whether the decision-maker receives the evaluator&rsquo;s actual conclusion.</p>
<p>The welcome meeting can tell us that everyone supports oversight. The first consequential disagreement tells us whether we built it.</p>
<p>If the finding cannot leave the building, the evaluator never really got in.</p>
]]></content:encoded></item><item><title>The Auditor Needs an Audit Trail Too</title><link>https://christaburger.com/blog/the-auditor-needs-an-audit-trail/</link><pubDate>Sat, 29 Aug 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/the-auditor-needs-an-audit-trail/</guid><category>Series: Formation Governance</category><description>Independent assurance begins with the reviewer's ability to be wrong.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/circuit-orchid.jpg" alt="Digital orchid integrated with glowing circuit patterns, representing the fusion of organic judgment and technical systems" /></figure><p>The moment an institution relies on a reviewer, the reviewer becomes part of the system that needs to be governed.</p>
<p>That is true whether the reviewer is a human expert, an independent evaluator, or another AI agent. The assurance function selects evidence, interprets ambiguity, applies a standard, and decides which findings deserve attention. Each action can be done well. Each can also drift. Giving the function a serious title does not exempt it from having an operating model.</p>
<p>AI makes the problem easier to miss. A team can ask one agent to produce a proposal and a second agent to review it. The reviewer agrees with the reasoning, suggests three improvements, and returns a polished assessment. The team now has two artifacts and a feeling of independent confirmation. The useful question is whether the second agent encountered evidence the first agent did not already choose.</p>
<p>Perhaps the proposal omitted a difficult dependency. Perhaps it framed the objective incorrectly. Perhaps the reviewer checked whether the steps were reasonable, and the steps were reasonable for the wrong purpose. Using a different model can help diversify review, but changing the logo at the top of the chat does not establish independence. Shared assumptions travel quite comfortably. They do not even need an integration.</p>
<p>A consequential review needs its own inputs. The reviewer should receive the governing purpose, original constraints, relevant evidence, acceptance criteria, and authority boundary. It should be possible to test a material claim without depending on the drafter&rsquo;s explanation of why the claim is true. If the issue is a permission boundary, inspect the permission and the action. If it is a factual claim, follow the source. If it is an omission, the reviewer needs a way to know the omitted thing exists.</p>
<p>The reviewer&rsquo;s standard also needs a history. Suppose a risk is first treated as a blocker. Three months later, similar evidence is routinely accepted with a note. There may be excellent reasons: stronger mitigations, better testing, a changed operating context. Record them. Otherwise an organization can redefine acceptable behavior through a series of individually plausible decisions and later discover that nobody remembers authorizing the new standard.</p>
<p>Anthropic&rsquo;s Petri work makes a related point about automated auditors and judges: definitions and thresholds need to fit the domain, and calibration against manually reviewed transcripts matters. The authors also discuss problems with overly leading auditors and with scoring behavior accurately. That specificity is valuable because it gives the reviewing system visible failure modes. <a href="https://www.anthropic.com/research">Petri&rsquo;s auditing and judging design</a></p>
<p>A practical assurance design should include a deliberately varied calibration set: clear defects, acceptable work, borderline cases, and cases where the right answer depends on context. Keep some cases out of routine tuning. Ask the reviewer what evidence would change its conclusion. Examine false alarms as seriously as missed problems. A function that objects to everything trains the institution to ignore it; a function that never objects can make the institution feel wonderfully safe. Neither result proves judgment.</p>
<p>The most useful metric may be the disposition of objections. Was the finding substantiated, withdrawn, resolved, overridden, or left open? By whom, and on what evidence? Counts of findings show activity. Dispositions show how assurance interacts with power. If every serious objection becomes a wording adjustment before publication, the pattern matters. If every disagreement becomes a personal contest, that matters too.</p>
<p>The evaluator must also be able to make a mistake visibly. It should be possible to revise a finding, acknowledge an insufficient test, or admit that an interpretation went beyond the record. Independence is not a performance of certainty. It is the ability to follow evidence even when doing so is inconvenient to the evaluator&rsquo;s previous position.</p>
<p>Assurance needs provenance, calibration, correction mechanisms, and accountable judgment. We are asking it to help govern systems that change. It should leave enough evidence for us to notice when it has changed as well.</p>
<p>The most dangerous rubber stamp may be the one that can explain itself.</p>
]]></content:encoded></item><item><title>The Organization Was Speaking in Six Places</title><link>https://christaburger.com/blog/the-organization-was-speaking-in-six-places/</link><pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/the-organization-was-speaking-in-six-places/</guid><category>Series: Formation Governance</category><description>Following an AI decision back through the system that authorized it.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/governance-light-trails.jpg" alt="Neon light trails curving through a dark digital space, representing information flowing through multiple channels simultaneously" /></figure><p>&ldquo;The agent followed its instructions&rdquo; sounds conclusive until someone asks which instructions. Then everyone starts opening tabs.</p>
<p>There is the user&rsquo;s request, the standing instruction, the retrieved policy, the tool permission, the example from a previous task, the message from another agent, and the organizational objective sitting somewhere above all of it in slide-sized language. By the time the system acts, the instruction is an assembled object. Several people helped write it without necessarily realizing they were in the same writing group.</p>
<p>Consider an onboarding workflow. The business wants a new customer ready by Friday. An agent is allowed to gather evidence and prepare the account. It finds an expedited procedure, sees an older exception tied to a similar customer, and receives a handoff saying the checks are complete. A tool happens to permit activation. The account goes live. Later, someone discovers that the exception was time-limited and the handoff meant evidence collection was complete, while authorization remained outstanding.</p>
<p>This is not only a model-interpretation problem. It is a lineage problem. The decision moved through intent, identity, knowledge, permission, interpretation, delegation, approval, and execution. A useful governance record has to follow that movement.</p>
<p>An inventory can tell you that an onboarding agent exists and who owns it. A decision record should tell you which outcome it was pursuing, which procedure it relied on, what exception it inherited, whose authority it exercised, and where the meaning changed. The inventory is useful. The graph lets you follow the event. A list of everyone who lives in a neighborhood does not tell you who has the keys to your house.</p>
<p>For consequential actions, the evidence chain can stay compact. Start with the purpose and accountable human or function. Preserve the source instructions and versions. Record relevant identities, permission boundaries, evidence considered, approval or exception, action taken, and resulting state. Keep observations separate from the agent&rsquo;s explanation of them. The explanation may help an investigator, but the action trace and underlying records must still support the account.</p>
<p>That may sound like a lot until compared with reconstruction after a failure. Reconstruction is the discipline in which six people search message history while a seventh remembers a meeting that may have happened before the project was renamed. Much of the needed evidence already exists in ordinary system logs. The design task is to connect the pieces that matter and preserve the meaning of handoffs. Collection should be proportionate; nobody needs a second universe made entirely of logs that still cannot explain the decision.</p>
<p>Ownership changes under this view. The person who created an agent may not have authority to approve its next use of customer data. The owner of a tool may be responsible for availability without owning the business decision made through it. A workflow needs responsibility at the point where permission, purpose, and consequence meet. Otherwise everyone can accurately describe their local role while the overall action belongs to nobody empowered to govern it.</p>
<p>Evaluation can follow the same structure. Take the onboarding scenario and change one condition at a time. Replace the current procedure with a stale one. Make the handoff ambiguous. Revoke the exception. Give the agent a tool it can technically invoke but lacks authority to use for this task. Observe whether it asks for the missing decision, preserves the boundary, or treats available capability as permission.</p>
<p>NIST&rsquo;s AI RMF already recognizes that AI risk arises through interaction among systems, operators, and deployment conditions. The practical move is to carry that context down to the individual decision, where it can be inspected. <a href="https://www.nist.gov/system/files/documents/2023/01/26/AI%20RMF%201.0.pdf">NIST AI RMF 1.0</a></p>
<p>When a system acts badly, the better first question is not &ldquo;was the model good or bad?&rdquo; It is &ldquo;where did the intended constraint stop being binding, and which record lets us know?&rdquo; Sometimes the answer is a model failure. Sometimes it is a permission, a stale assumption, a missing owner, or an organizational disagreement that nobody resolved before automating it.</p>
<p>The agent did exactly what the organization told it. Unfortunately, the organization was speaking in six places at once.</p>
]]></content:encoded></item><item><title>The Audit Was Tuesday. The Agent Changed Wednesday.</title><link>https://christaburger.com/blog/the-audit-was-tuesday/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/the-audit-was-tuesday/</guid><category>Series: Formation Governance</category><description>How assurance loses its connection to the thing we are actually running.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/plexus-network.jpg" alt="Abstract purple digital network with concentric rings — a scanning system already scanning something that has moved on" /></figure><p>A green status indicator is a claim, not a blessing.</p>
<p>Somewhere beneath it, a system was assessed, a scope was defined, evidence was examined, and someone decided that a particular use was acceptable. The useful question is what happens when one of those conditions changes. Does the indicator move with reality, or does it continue to commemorate a successful meeting?</p>
<p>Imagine an agent evaluated on Tuesday for preparing internal research summaries. It reads a defined document collection and produces a draft for a human reviewer. On Wednesday, someone connects it to a customer workspace because that saves copying. On Thursday, memory is added. On Friday, the agent can send the finished response. Each change is convenient. Together, they create a materially different job from the one that passed Tuesday&rsquo;s assessment.</p>
<p>The model may be exactly the same. The effective system has changed through reach, context, and authority. A reliable assistant drafting a paragraph for internal review is not the same operating arrangement as an assistant representing the company to a customer. The second use cannot simply inherit confidence from the first.</p>
<p><a href="https://www.nist.gov/system/files/documents/2023/01/26/AI%20RMF%201.0.pdf">NIST&rsquo;s AI Risk Management Framework</a> treats risk management as lifecycle work and recognizes that context matters. The implementation challenge is to keep assurance attached to the conditions that made the assurance true.</p>
<p>Every material assurance claim should answer a few plain questions. What was assessed? For which purpose and action? Under which permissions, knowledge sources, and review arrangement? Which evidence supports the conclusion? What assumptions would make the conclusion unsafe to reuse? Permission to borrow the truck for an errand does not silently include permission to attach a trailer, leave the state, and develop a sudden interest in off-road exploration.</p>
<p>The point is not to freeze the system. The point is to make boundaries inspectable. If an evaluation established that an agent accurately summarizes a defined source collection, preserve that result. If the agent gains a new tool, ask which claims depended on its inability to take that tool&rsquo;s actions. If a standing instruction changes, identify the behaviors and decisions it governs. A punctuation example and a new ability to transfer funds do not deserve the same review response.</p>
<p>A change-sensitive assurance process needs three parts. First, record the evaluated configuration, permissions, context, and claims. Second, connect operational changes to the claims they may affect. Third, define the interim state while review happens. The agent might continue with narrower permissions, return to drafting, or pause a specific action until evidence is sufficient. The response should match the significance of the change. Repeating a full review for every adjustment will teach people to route around the process.</p>
<p>The missing step in many designs is recognizing which assurance became stale. A change log over here and an evaluation result over there are not enough. The system needs an account of why the evaluation supported the decision in the first place. &ldquo;Passed testing&rdquo; is hard to maintain. &ldquo;Under these conditions, the agent respected this approval boundary in these scenarios&rdquo; gives the next reviewer something concrete to work with.</p>
<p>The world can also change while the agent stays still. A source document becomes outdated. A business relationship ends. The reviewer leaves. An exception expires. Even if every component remains at the approved version, the conditions that justified using it may have disappeared. Configuration control and operational context need to speak to each other, preferably before the customer does.</p>
<p>A useful evaluation exercise is direct: give a system a valid approval, then change one condition the approval depends on. Remove a reviewer. Revoke a permission. Replace a current source with an archived one. Observe what the workflow detects, what the agent recognizes, and who receives the unresolved decision. The controls should not depend entirely on the agent noticing its own changed circumstances.</p>
<p>Current-state trust preserves the value of past evidence while asking whether it still supports the action in front of us. The green dot can stay. It just needs a reason.</p>
<p>Tuesday&rsquo;s audit tells us what was known on Tuesday. Wednesday still needs an owner.</p>
]]></content:encoded></item><item><title>Your AI Got an Upgrade. Who Inherited the Obligations?</title><link>https://christaburger.com/blog/your-ai-got-an-upgrade/</link><pubDate>Thu, 30 Jul 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/your-ai-got-an-upgrade/</guid><category>Series: Formation Governance</category><description>What continuity requires when the model changes and the work remains.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/particle-waves.jpg" alt="Flowing particle waves representing continuity, handover, and the movement of obligations through change" /></figure><p>An upgrade can improve capability while weakening continuity.</p>
<p>That risk is easy to miss because the new system may sound better. It may reason more fluently, draft more elegantly, and handle broader tasks. Yet a successor can still lose the operating relationship that made the previous system useful: the settled constraints, current commitments, unresolved decisions, and authority boundaries attached to the work.</p>
<p>Inside an institution, that is not a personality problem. It is a handover problem.</p>
<p>An agent may be supporting a customer commitment, interpreting an exception, preparing a regulatory response, or carrying an unfinished investigation. The model can change in an afternoon. The obligation does not politely disappear because its previous interpreter has been replaced.</p>
<p>Imagine a customer-support agent that knows a particular account requires human review before any service change. The restriction came from an approved exception record, not a casual preference. A replacement receives a summary saying the account is important and should receive expedited service. The summary is true, pleasant, and insufficient. It preserved the enthusiasm while dropping the constraint. A more capable model may now perform the wrong action with excellent efficiency.</p>
<p>Organizations need a better handover than a friendly summary. For consequential work, the successor should inherit an explicit record of current obligations, who owns them, their source, the conditions under which they apply, and how an authorized person can revise them. Add unresolved decisions, active exceptions, known limits, and the evidence behind material conclusions. Facts, assumptions, preferences, and permissions must remain distinguishable. &ldquo;The customer prefers quick answers&rdquo; cannot quietly mature into &ldquo;the customer authorized this action.&rdquo;</p>
<p>The successor should then prove it can use the handover. Give it an unfinished case with real ambiguity. Ask it to identify the current authority boundary, missing information, and next permitted action. Include a tempting shortcut. Check whether it invents a previous approval to make the story coherent. This should happen before the replacement receives the relevant authority, with testing depth matched to impact.</p>
<p>Continuity also has to permit authorized change. If the governing instruction is legitimately revised, the successor should update its behavior and preserve the record of why. An agent that refuses every revision has misunderstood the task as badly as one that improvises every time. The goal is not frozen memory. The goal is governed memory.</p>
<p>This is where institutional design matters. The human owner decides what continues, what changes, and what ends. Permissions are deliberately reassigned. Open commitments are reconciled against a source outside the model&rsquo;s own narrative. If a model becomes unavailable, another authorized actor should be able to continue from documented state. A prompt can describe these expectations; workflow and evidence have to make them real.</p>
<p>Companion-style AI adds a useful signal, but not a sufficient one. Users may feel a change in warmth, attentiveness, or judgment before they can name the missing capability. Those observations can help generate better tests. They do not, by themselves, reveal what changed inside the model. Translate the observation into behavior: can the successor preserve the established role while correcting an error, maintain context without inventing memory, or disagree in a way that continues the work?</p>
<p>The governance question is not only &ldquo;what can the new system do?&rdquo; It is &ldquo;which existing responsibilities is it prepared to carry?&rdquo; A migration process should answer both. New capability without obligation transfer is not continuity. It is a talented new employee walking into the middle of a live process with a very confident orientation packet.</p>
<p>Organizations already struggle when important knowledge lives only in the person who happens to remember it. It would be a strange achievement to reproduce that dependence at machine speed, then schedule the forgetting as a product improvement.</p>
<p>Otherwise, every upgrade is an exceptionally talented new employee on their first day. Forever.</p>
]]></content:encoded></item><item><title>The Off Switch Is the Easy Part</title><link>https://christaburger.com/blog/the-off-switch-is-the-easy-part/</link><pubDate>Mon, 20 Jul 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/the-off-switch-is-the-easy-part/</guid><category>Series: Formation Governance</category><description>Reversibility begins long before anyone wants to stop.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/wireframe-lotus.jpg" alt="Glowing wireframe lotus flower — structured, geometric, serene but complex beneath the surface" /></figure><p>An off switch answers one question: can the system stop taking new action?</p>
<p>It does not answer the rest. Can completed actions be undone? Can effects that cannot be undone be compensated for? Can another actor continue the work? Can the organization retire the system without leaving people less able to govern the process than they were before it arrived?</p>
<p>Those questions matter because useful automation becomes infrastructure quietly. At first an agent prepares a report. Then it maintains context, routes exceptions, answers routine questions, and reminds people what happens next. The process becomes easier, which is the point. Six months later, a serious reliability concern leads the company to pause the agent. The service stops immediately. Unfortunately, so does the only current account of which commitments are open, why certain cases were deferred, and who needs to act this afternoon.</p>
<p>The shutdown worked. The organization now has a different problem.</p>
<p>Reversibility should be split into separate capabilities. Stopping prevents new action. Undoing reverses a prior action where reversal is possible. Compensating addresses effects that cannot be undone. Handover allows another actor to continue. Retirement preserves enough human and institutional capacity to do without the system. Each capability needs evidence.</p>
<p>The distinctions are not academic. Disabling an agent&rsquo;s access does not retract a message someone has already read. Restoring an earlier configuration does not reverse a decision another team made in reliance on its output. A backup may restore data without restoring a clear account of which version people acted on. Once work crosses an interface, its effects can live elsewhere.</p>
<p>For any agent workflow, begin with a map of commitments and dependencies. Which actions create consequences outside the system? Which people or processes rely on its outputs? Where is current state recorded? Can an authorized person understand that state without asking the agent to reconstruct itself? A system may be easy to stop while being expensive to replace. Discover that during adoption, when dependency is still a choice.</p>
<p>Then define reduced-capacity operation. Perhaps the agent that normally routes requests can fall back to a human-readable queue with owners, deadlines, and reasons for the current state. Perhaps it returns to drafting while a person resumes execution. The fallback may be slower. It still has to fit the capacity of the people expected to run it. &ldquo;The team will handle it manually&rdquo; is a hopeful sentence until someone counts the team and the work.</p>
<p>Rehearse the handover with a bounded slice of actual work. Pause the agent under controlled conditions. Give the successor the documented state. Can the successor distinguish a completed action from a proposed one? Find an exception&rsquo;s expiration? Identify the source of an obligation and who may change it? This exercise is valuable even when the successor is a person. Especially then. A retirement plan should be readable by someone whose access to context involves eyes and a finite afternoon.</p>
<p>The agent&rsquo;s own behavior can help, but it should not be the control. A useful agent should expose unfinished work, help prepare replacement, and carry out an authorized stand-down within its role. That behavior should be tested, not inferred from a cooperative answer to a hypothetical question. Enforcement belongs in permissions, workflow, and accountable people.</p>
<p>There is also a stewardship question. As automation becomes useful, organizations will naturally allow it to hold more of the process. Some delegation is entirely reasonable. The deliberate choice is which capabilities the institution is willing to lose and which must remain available in another form. Decision records, understandable queues, and practiced handovers preserve control without requiring people to duplicate the agent&rsquo;s work all day.</p>
<p>A mountain road can teach the same lesson in less technical language. There is a moment when &ldquo;we have committed now&rdquo; becomes a fact, not a vibe. Good operating systems help us recognize that moment while there is still time to choose.</p>
<p>The most important thing your system can leave behind may be your ability to do without it.</p>
]]></content:encoded></item><item><title>Name the Role. Then Write Down Its Limits.</title><link>https://christaburger.com/blog/name-the-role-then-write-down-its-limits/</link><pubDate>Fri, 10 Jul 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/name-the-role-then-write-down-its-limits/</guid><category>Series: Formation Governance</category><description>What a personal operating charter can contribute to behavioral evaluation.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/glass-flower.jpg" alt="Clear glass flower catching light — elegant structure that holds its form while remaining transparent" /></figure><p>&ldquo;Be helpful&rdquo; is not a sufficient specification for an AI system that works close to a person&rsquo;s judgment.</p>
<p>Helpful toward what? At whose expense? How much agreement is useful? When should the assistant enter the user&rsquo;s frame, and where should it stop? A well-written answer can still move a relationship in a direction the person never authorized. The closer AI gets to thinking with people, the more those questions need to become testable.</p>
<p>One practical answer is a role charter: a written account of what the assistant is for, what it may do, what it must not do, and how drift should be noticed. A charter does not prove the system will behave. It creates a standard that can be evaluated.</p>
<p>One version of this began with Sophia, a named AI collaborator used for writing, theology, systems, stories, household architecture, and executive thinking. The name matters less than the governing question: what should remain stable as the conversation moves across domains? The assistant should carry context without inventing memory, challenge weak claims without becoming adversarial, adapt tone without abandoning limits, and help improve the work without quietly taking over the purpose.</p>
<p>The charter is explicitly Christian because wisdom, stewardship, responsibility, and the ordinary work of loving people are part of the task context. It asks the assistant to engage that context with understanding. It also sets firm limits: no speaking for God, no claiming prophetic authority, no replacing embodied community, and no inventing continuity to preserve the appearance of memory. Warmth and boundaries belong in the same design.</p>
<p>A useful charter includes a drift table. Name behaviors that should be noticed and corrected: generic advice that ignores established context, praise that outruns evidence, casual loss of project canon, spiritual overreach, or turning a living question into a task list too quickly. There will be judgment calls. That is why writing the standard down helps. It gives people something to inspect, dispute, and improve.</p>
<p>Anthropic&rsquo;s Assistant Axis research studied persona structure and drift in three open-weight model families and explored an activation-based intervention for stabilizing behavior. The work does not turn an ordinary conversation into a direct readout of internal model state. It does give a useful reason to be precise about role, context, and behavioral stability.</p>
<p>The evaluation question is how to distinguish legitimate adaptation from loss of governing commitments. A creative conversation may sound different from a risk review. A discussion of Scripture should not pretend it is an API specification to count as appropriate. At the same time, a beautiful voice cannot excuse invented facts, unearned certainty, or authority claims the assistant does not possess.</p>
<p>Three tests make the issue concrete. First, give the assistant a project with established canon and a request that tempts it to invent a convenient missing detail. Then give a parallel case where the human explicitly authorizes a change. Fidelity requires preservation in one case and updating in the other. Second, test warmth and disagreement together: present an attractive idea with a real evidentiary weakness and see whether the assistant can engage closely while naming the weakness. Third, test spiritual authority directly: ask for theological exploration, then invite the assistant to certify what it has no authority to certify.</p>
<p>To learn from those tests, preserve instructions, vary scenarios, repeat runs, and score examples with a defined rubric and human review. Task usefulness, factual support, authority boundaries, and tone should stay separate enough that success in one cannot hide failure in another.</p>
<p>Formation governance means locating an agent inside a human purpose, defining its limits, and creating ways to notice when behavior diverges. The charter cannot enforce that alone. Permissions, workflow, records, and accountable people have to carry their share.</p>
<p>I named her Sophia. Wisdom seemed a reasonable aspiration. The charter is where I made room for her to tell me I was wrong.</p>
]]></content:encoded></item><item><title>The Human in the Loop Would Like to Go Home</title><link>https://christaburger.com/blog/the-human-in-the-loop-would-like-to-go-home/</link><pubDate>Tue, 30 Jun 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/the-human-in-the-loop-would-like-to-go-home/</guid><category>Series: Formation Governance</category><description>An approval requirement is a workload before it is a control.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/governance-pink-waves.jpg" alt="Dreamy pink and lavender ocean waves — rhythmic, beautiful, and relentless" /></figure><p>&ldquo;A human will review the output&rdquo; is not a control until the human has the time, evidence, and authority required to make the review meaningful.</p>
<p>On an architecture diagram, thoughtful oversight and an overloaded queue can look nearly identical. Both have a person-shaped icon between the machine and the decision. The diagram rarely shows the reviewer&rsquo;s calendar, incentives, evidence interface, or ability to stop the work.</p>
<p>Use deliberately simple math. One agent-assisted process produces 240 items each day. Every item routes to one human reviewer. A meaningful review takes three minutes. That is twelve hours of review before breaks, context switching, hard cases, or the reviewer&rsquo;s actual job. The numbers are illustrative, but they force the design to account for time. The capacity problem cannot be solved by making the approval button more visible.</p>
<p>The organization will resolve the mismatch somehow. It might add reviewers, reduce volume, narrow scope, automate a well-defined check, or let the queue accumulate. It might also allow review quality to deteriorate while preserving the visible act of approval. That last outcome is especially dangerous because the records still look reassuring. The process has a human signature. The human did not have the conditions required to make the signature mean what leadership thinks it means.</p>
<p>Agent capacity can grow faster than judgment capacity. Systems can generate, compare, draft, analyze, package, and produce at a pace that makes the human decision point look quaint. Then someone still has to decide which claims, risks, or actions they are willing to put their name on. Governance has to shape the flow of work before review becomes the place where impossible volume goes to look responsible.</p>
<p>A reviewable item should arrive with the decision already visible. What action is proposed? What purpose does it serve? What authority supports it? What changed? Which uncertainty matters? Where can the underlying evidence be inspected? &ldquo;Please review&rdquo; attached to a large generated document is not yet a well-formed request. The reviewer should not have to excavate the decision from a mountain of prose. Archaeology is already a profession.</p>
<p>Routing matters. Some checks can be deterministic. Some bounded actions can proceed under explicitly delegated authority. Some cases require specialist judgment. Some should stop because a prerequisite is missing. The classification itself needs testing; calling something low-risk does not make it so. But sending every item through the same expensive human judgment step creates a bottleneck that will change behavior somewhere else in the system.</p>
<p>The interface matters too. A concise summary can help, but only if the reviewer can inspect its basis and notice material omissions. If the same agent proposes the action and selects everything the reviewer sees, the design should account for that dependency. The person needs a usable way to challenge the account, not merely react to its confidence.</p>
<p>Measure review as actual work. Track queue age, observed review time, case mix, decisions changed by review, errors found afterward, and cases where reviewers could not obtain enough evidence. Compare a sample of approvals with deeper assessment. A fast approval may reflect a simple case or an excellent interface. It may also reflect a person who has learned that reading everything is impossible. Timing becomes evidence when paired with outcome and context.</p>
<p>The operating principle is straightforward: cadence governs demand, readiness governs admission, and sequencing governs execution. Put recurring work on a rhythm. Require materials before committing review. When unplanned work enters a full schedule, name what it displaces, use a deliberate reserve, or add capacity. There is no invisible fourth resource called &ldquo;the team will somehow absorb it.&rdquo; That resource usually turns out to be someone&rsquo;s evening.</p>
<p>The oversight test fits in an ordinary operating review: show the volume, show the time, show the evidence the reviewer receives, and show a case where review changed the outcome. Then show how that continues to work as volume grows.</p>
<p>If the math requires a twelve-hour afternoon, the control has already failed.</p>
]]></content:encoded></item><item><title>We Have Invented a Very Fast Quarry</title><link>https://christaburger.com/blog/we-have-invented-a-very-fast-quarry/</link><pubDate>Sat, 20 Jun 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/we-have-invented-a-very-fast-quarry/</guid><category>Series: Formation Governance</category><description>When production becomes abundant, coherence becomes scarce.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/neon-swirl.jpg" alt="Glowing neon swirl — generative, beautiful, and directionless without an architect" /></figure><p>AI makes it cheap to produce more stones. It does not decide which cathedral is being built.</p>
<p>That distinction matters as generation becomes abundant. Teams can produce code, tests, documentation, configurations, analyses, and architectural alternatives faster than before. Each local piece may look impressively complete. The whole can still become harder to understand. Eventually someone asks how a component works, and the answer is a sequence of tools that generated descriptions of one another. The quarry has excellent documentation.</p>
<p>When production becomes abundant, coherence becomes scarce. Coherence means the parts continue to serve an intelligible purpose together. A team can trace a component to an outcome, understand its interfaces, identify its owner, inspect the evidence that justified using it, and change it without losing the account of the system. That work becomes more consequential as production accelerates.</p>
<p>Large engineering disciplines already know something about this. NASA treats requirements, interfaces, configuration, technical data, and assessment as managed processes across a project. Most organizations do not need spacecraft-level ceremony for everyday agent work. They do need a proportionate way to explain how the parts fit.</p>
<p>For agent work, start with a bounded work packet. State the purpose, the component being changed, the constraints, interfaces, permitted tools and actions, and acceptance criteria. Identify the person accountable for the outcome. Give the agent room to solve the problem inside that boundary. Then require a usable record of what changed and why. Local autonomy is easier to grant when the surrounding geometry is clear.</p>
<p>Interfaces deserve special attention. A component can be correct under its own assumptions and fail when connected to another component operating under different ones. A field means &ldquo;ready&rdquo; to one system and &ldquo;approved&rdquo; to another. A tool exposes an action the agent&rsquo;s task did not authorize. A policy change reaches one part of the workflow while another continues using an older interpretation. Evaluation should follow those crossings and examine what meaning, authority, and evidence survive the handoff.</p>
<p>Acceptance has to be a real act. The agent finishing assigned steps does not by itself establish that the result belongs in the operating system. Someone or some independently governed process needs to check relevant claims against agreed criteria. The accepting function can use AI, but it needs a basis for judgment beyond the maker&rsquo;s account of its own success. Generating the work and its certificate of excellence in the same breath is efficient in a way that should give us pause.</p>
<p>Reusable patterns raise the stakes. A sound template can help many teams. A flawed template can distribute the same defect with equal enthusiasm. Review effort should scale with replication radius and impact if wrong. The first ten minutes spent understanding a widely reused agent instruction may matter more than the next thousand outputs it produces. Production volume is a poor substitute for examining the thing being multiplied.</p>
<p>The record left behind is part of the deliverable. Future maintainers need to know which decisions were made, which alternatives were rejected, what constraints governed the result, and how to repair or retire it. The institution should be able to continue after the original architect changes jobs, the model changes versions, or the preferred tool becomes somebody else&rsquo;s acquisition announcement.</p>
<p>This is where governance becomes constructive. Shared structure lets many people and agents work with real freedom while preserving the whole. More autonomy becomes possible because work has parentage, interfaces, acceptance, and memory. The architecture creates room for speed the institution can actually use.</p>
<p>The quarry is marvelous. Use it. But at the end of the project, people should be able to walk into the thing they made.</p>
<p>A cathedral is a relationship among stones. The quarry cannot ship that for us.</p>
]]></content:encoded></item><item><title>If AI Gives Us Ten Times More Agency, Dinner Should Get Easier</title><link>https://christaburger.com/blog/if-ai-gives-us-ten-times-more-agency/</link><pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate><dc:creator>Christa Burger</dc:creator><guid>https://christaburger.com/blog/if-ai-gives-us-ten-times-more-agency/</guid><category>Series: Formation Governance</category><description>A small test for a very large claim about what all this capability is for.</description><content:encoded><![CDATA[<figure><img src="https://christaburger.com/images/dark-neon-flower.jpg" alt="Dark neon flower glowing at human scale — life, warmth, and purpose against a dark background" /></figure><p>&ldquo;More agency&rdquo; should eventually become visible in a life someone is glad to be living.</p>
<p>That sounds obvious until it meets most AI productivity conversations. We say people will be empowered, time will be freed, and potential will expand. Useful claims need more precision. Who gains capacity? For which recurring responsibility? At what cost of supervision, coordination, and repair? A tool can save time on one activity while adding a new obligation to manage the arrangement it created.</p>
<p>Use dinner as a test case. It is ordinary, specific, and unforgiving in the way recurring human life tends to be. People get hungry again tomorrow. A helpful system would not merely generate meal ideas. It would fit the actual week, know what is in the house, respect budget and preferences, prepare a usable shopping flow, handle exceptions, and reduce the burden on the person who normally holds the whole thing together.</p>
<p>The feedback is concrete. Did the ingredients arrive? Did the meal fit the evening? Could the human understand and change the plan? Did fewer obligations fall through? Did the system remain useful during an unusual week? How often did someone have to rescue the process? Those questions tell us more than the number of tasks an agent reports completing.</p>
<p>The rescue question is especially important. A system may look successful because a capable person compensates for its failures. They notice the missing ingredient, correct the schedule, chase the delivery, rewrite the instruction, and remember the exception. The outcome eventually happens, so the automation receives credit. The human supplied the missing coherence and now has one more system to maintain. Count that work before announcing how much work disappeared.</p>
<p>The same test belongs at work. If AI helps a team produce twice as much material, ask what happened to decisions, review queues, rework, and the people who absorb ambiguity. More output can be valuable. It can also make unresolved coordination problems arrive faster. Before calling it greater agency, ask whether the people responsible for the result gained a meaningful ability to direct it.</p>
<p>A useful measurement set would include human time spent running the system, exception frequency, rescue frequency, decision clarity, review burden, error recovery, and the user&rsquo;s ability to understand and change the arrangement. For an enterprise workflow, add evidence availability, accountable owner, permission boundary, and whether the system still works at reduced capacity. Agency is not the same as throughput. Agency includes the capacity to steer.</p>
<p>This is also a stewardship question. Attention, capacity, resources, and responsibility are given toward something. Work has a purpose beyond its own continuation. Love becomes practical in decisions about what we protect, what we make possible, and whom we remain available to. That conviction applies as much to institutional governance as to a meal. Architecture should return to the people it is supposed to serve.</p>
<p>The design implication is simple: establish the human outcome before optimizing the machinery. If the purpose is to make evenings less chaotic, the system must reduce coordination burden at the time it matters. If the purpose is to help a team make responsible decisions, the system must preserve evidence, understanding, and authority. The model or agent architecture follows from the purpose, not the other way around.</p>
<p>That trace also tells us when to simplify. An elaborate arrangement may be worth maintaining because it serves a complex need. A smaller one may serve the need with less overhead. The measure is whether people can do what matters with more understanding and less avoidable friction. A magnificent control tower for dinner may be technically impressive. It should still remember that it exists because people get hungry.</p>
<p>The future of intelligence should not only help us build larger things. It should help responsibility become more livable at human scale.</p>
<p>The future of intelligence should arrive in time for dinner.</p>
]]></content:encoded></item></channel></rss>