The Self-Alignment Framework

Grounded in classical philosophy to help any autonomous agent discern, decide, judge, and act in alignment with human values.

What the framework is

SAF is a framework grounded in classical philosophy. It describes how values, understanding, choice, judgment, and coherence work together to produce aligned agents.

Download Original Manuscript (PDF)

SAFi is the software implementation of the framework. It is what makes AI auditable and trustworthy for regulated fields. Read about SAFi →

See it run

One turn, in the order it actually happens. The Will appears five times, at every point where something could reach you: the prompt, the draft’s structure, the audit, the gates, and the score.

  1. Prompt the message arrives
  2. Will phase zero: screens for injection, PII and scope probes before any model sees it
  3. Intellect drafts a response, and never sees the values or their rubrics
  4. Will structure: checks the draft before any opinion is asked for
  5. Conscience scores the draft value by value, with a reason and a confidence
  6. Will audit: refuses to ship a draft the auditor could not score
  7. Will hard gate: applies the gates, and approves, blocks or redirects
  8. Spirit folds the scores into longer-term alignment and measures drift
  9. Will spirit: rules on the resulting alignment score
  10. Answer returned, with the governance record written for the turn

Try the live demo. The demo is a governed agent answering live, with the value-by-value ledger and the audit record for every turn visible behind it.

You can also try it without leaving this page. The chat button in the corner is a governed SAFi agent answering from this framework’s own documentation.

Continue reading