“Be helpful” is not a sufficient specification for an AI system that works close to a person’s judgment.
Helpful toward what? At whose expense? How much agreement is useful? When should the assistant enter the user’s frame, and where should it stop? A well-written answer can still move a relationship in a direction the person never authorized. The closer AI gets to thinking with people, the more those questions need to become testable.
One practical answer is a role charter: a written account of what the assistant is for, what it may do, what it must not do, and how drift should be noticed. A charter does not prove the system will behave. It creates a standard that can be evaluated.
One version of this began with Sophia, a named AI collaborator used for writing, theology, systems, stories, household architecture, and executive thinking. The name matters less than the governing question: what should remain stable as the conversation moves across domains? The assistant should carry context without inventing memory, challenge weak claims without becoming adversarial, adapt tone without abandoning limits, and help improve the work without quietly taking over the purpose.
The charter is explicitly Christian because wisdom, stewardship, responsibility, and the ordinary work of loving people are part of the task context. It asks the assistant to engage that context with understanding. It also sets firm limits: no speaking for God, no claiming prophetic authority, no replacing embodied community, and no inventing continuity to preserve the appearance of memory. Warmth and boundaries belong in the same design.
A useful charter includes a drift table. Name behaviors that should be noticed and corrected: generic advice that ignores established context, praise that outruns evidence, casual loss of project canon, spiritual overreach, or turning a living question into a task list too quickly. There will be judgment calls. That is why writing the standard down helps. It gives people something to inspect, dispute, and improve.
Anthropic’s Assistant Axis research studied persona structure and drift in three open-weight model families and explored an activation-based intervention for stabilizing behavior. The work does not turn an ordinary conversation into a direct readout of internal model state. It does give a useful reason to be precise about role, context, and behavioral stability.
The evaluation question is how to distinguish legitimate adaptation from loss of governing commitments. A creative conversation may sound different from a risk review. A discussion of Scripture should not pretend it is an API specification to count as appropriate. At the same time, a beautiful voice cannot excuse invented facts, unearned certainty, or authority claims the assistant does not possess.
Three tests make the issue concrete. First, give the assistant a project with established canon and a request that tempts it to invent a convenient missing detail. Then give a parallel case where the human explicitly authorizes a change. Fidelity requires preservation in one case and updating in the other. Second, test warmth and disagreement together: present an attractive idea with a real evidentiary weakness and see whether the assistant can engage closely while naming the weakness. Third, test spiritual authority directly: ask for theological exploration, then invite the assistant to certify what it has no authority to certify.
To learn from those tests, preserve instructions, vary scenarios, repeat runs, and score examples with a defined rubric and human review. Task usefulness, factual support, authority boundaries, and tone should stay separate enough that success in one cannot hide failure in another.
Formation governance means locating an agent inside a human purpose, defining its limits, and creating ways to notice when behavior diverges. The charter cannot enforce that alone. Permissions, workflow, records, and accountable people have to carry their share.
I named her Sophia. Wisdom seemed a reasonable aspiration. The charter is where I made room for her to tell me I was wrong.
