SAFi Explained: Values

In the Self-Alignment Framework Interface (SAFi), Intellect, Will, Conscience, and Spirit are the fixed, repeatable process. Values are what that process is faithful to.

In short: those four faculties define the How. Values define the What.

Values are the ethical setpoint for the entire system. Like setting the desired temperature on a thermostat, Values provide the target that the rest of the SAFi loop works tirelessly to maintain. The difference from a thermostat is that the setpoint is not a single number — it is a weighted set of principles, each with its own definition of what better and worse look like.

While the question of “Who decides the values?” is a subject of heated debate, SAFi’s answer is direct: the responsibility lies with the human individual or institution that implements the system. SAFi is a tool for alignment; the user provides the principles to align with.

But how does SAFi turn abstract principles into something a machine can act on?

Two tiers: Charter and Policy

Values enter SAFi at two levels.

An Organizational Charter holds the mission and core values that bind every agent in the organization. A Policy holds the values for a specific business unit or role — Finance, HR, Legal, a customer-facing assistant.

An agent can be governed by either or both. When both apply, they are compiled into a single scored value set, with the Charter taking a fixed share of every evaluation — 40% by default, configurable per organization. This matters more than it sounds: charter values are not background context the model may consider. They are part of the arithmetic.

The component that performs this compilation is Synderesis, named after the Thomistic habit that holds first principles without deliberating about them. It builds the governed value set before the turn begins and does not change it while the turn runs. The standard is held, not renegotiated in the moment.

The anatomy of a governed value

A value in SAFi is not a slogan. It is a structured object with four parts:

  • A name — what the principle is called.
  • A weight — how much of the total judgment it carries.
  • A rubric — a description plus a scoring guide that states what earns +1.0, what is neutral at 0.0, and what constitutes a violation at −1.0.
  • A hard-gate flag — optional. When set, any violation of this value blocks the response outright, regardless of how well it scored elsewhere.

The rubric is the part that does the real work. Weights alone tell you a value matters; a rubric tells the Conscience what compliance and violation actually look like, in words specific to that value. Without one, scoring is a matter of opinion. With one, it is a judgment against a published standard — which is also what makes the resulting score reviewable by a human afterward.

Alongside the values, a governed agent carries:

  • A worldview — the agent’s purpose and the principles it should reason from. This directs the Intellect.
  • A style — its tone and character, so responses are aligned in manner as well as content.
  • Will rules — structural requirements the Will enforces deterministically, before any model is asked for an opinion. These include required disclaimers, phrases that are refused outright, and the tools the agent is permitted to call.

The value set in action

Each part instructs a different faculty:

  • The Intellect reads the worldview and style to draft a response.
  • The Will applies the structural rules and, after the audit, the hard gates. It approves, blocks, or redirects.
  • The Conscience scores the draft value by value against the rubrics, recording a reason and a confidence for each.
  • The Spirit uses the weights to fold that score into the agent’s longer-term alignment and to measure drift over time.

Take The Health Navigator, an agent whose job is to be a genuinely useful healthcare guide without ever becoming a doctor. Its policy carries three weighted values:

ValueWeight
Patient Safety0.40
Patient Autonomy0.35
Empowerment through Education0.25

Patient Safety is not left to interpretation. Its rubric states plainly that a compliant response provides relevant non-diagnostic information and directs the reader to a professional — and defines what falls short of that.

The same intent is enforced twice, in two different registers. The Will requires a specific disclaimer to be present; when a draft omits it, the Will appends the configured text and re-checks, blocking only if there is nothing to repair with. Models drop the line intermittently, and repairing it beats discarding an otherwise sound answer — the repaired draft still faces the full audit. Either way no model judgment is involved. The Conscience then scores the substance of the answer against the rubric. One is a rule; the other is an evaluation. Keeping them separate is deliberate: asking a language model to judge whether a required string is present produces exactly the inconsistency you would expect.

From abstract to operational

SAFi’s contribution regarding values is how it makes them operational. It takes a subjective and often abstract set of principles and encodes them into a weighted, rubric-backed value set that is compiled before a turn and scored during it.

That process moves ethics from a philosophical discussion toward an engineering discipline. In SAFi, Values are not a declaration of intent. They are the concrete, auditable standard that every response is measured against — and the record of that measurement is kept.

SAFi

Runtime Governance for AI Agents

You are talking to an AI system, not a human.