← Standards alignment

Recommended alignment · NIST AI RMF 1.0 · AI Risk Management Framework

How we recommend aligning your private AI with the NIST AI Risk Management Framework.

Where 800-207 governs who can connect, the AI RMF governs what the AI may do. Its Core has four functions — Govern, Map, Measure, Manage — and nineteen categories. For each one, here's what it requires, the approach we recommend, the configuration that implements it, and how you confirm it's working. This is a recommended reference — the exact configs are finalized against your deployment.

Source: NIST AI Risk Management Framework (AI RMF 1.0) ↗ · category structure per the NIST AIRC Core reference ↗ · requirements paraphrased; read the original for the authoritative text.

How to read this. "Requirement" is our plain-language reading of the category — not a quote. "Configuration" is the recommended recipe on an OpenZiti + local-model appliance; yours may differ. Four categories are marked "organizational practice" — they are human commitments no configuration can honestly claim, so we say what the deployment gives you to do the work with instead. Nothing here claims certification.
Govern 6Map 5Measure 4Manage 4 All 19 categories · complete

Govern

The culture the other three functions run inside — policy, accountability, and the supply chain the AI is built from. 6 categories

Govern 1Requirement
Policies and processes for mapping, measuring and managing AI risk are in place, transparent, and actually implemented — not just written down.
What we recommend
Make the policy the same artifact as the enforcement. One policy file drives the guardrails and generates the human-readable AI policy document — so the written policy and the running behaviour cannot drift apart. A policy that lives only in a PDF is the failure mode this category exists to catch.
The configuration
# The policy file IS the enforcement — the document is generated from it
policy:
  version: 1.4
  owner: "COO"
  review_cycle: quarterly
  renders_to: /var/hermes/ai-policy.md   # the readable version
guardrails:
  mode: enforced          # enforced | advisory
  approvals: sensitive     # every | sensitive | none
See it work
$ hermes policy show
policy v1.4   owner=COO   next review 2026-10-01
document ..... generated from policy v1.4
drift ........ NONE   # written policy == enforced policy
Govern 2Requirement
Accountability structures are in place — roles, responsibilities and lines of authority for AI risk are documented and staffed.
What we recommend
Give every role a named human owner in the config, and let no capability exist without one. Roles are the same objects that scope what the AI may do, so "who is accountable" and "what is permitted" are one list, not two — and an unowned role is a configuration error, not an oversight.
The configuration
# Every role is owned by a person, and scopes what the AI may do
roles:
  - name: legal-reader
    owner: "j.doe@firm.example"     # accountable human
    may: [read_file, search_docs]
  - name: ops-editor
    owner: "m.ruiz@firm.example"
    may: [read_file, write_file]
    approvals: every              # stricter than the default
require_owner: true                 # unowned role = refuses to load
See it work
$ hermes roles list
legal-reader  owner=j.doe    may=read_file,search_docs
ops-editor    owner=m.ruiz   may=read_file,write_file   approvals=every
# adding a role with no owner:
ERROR role 'temp-bot' has no owner — refused to load
Govern 3Requirement · organizational practice
Workforce diversity, equity, inclusion and accessibility are prioritized in how AI risk is mapped, measured and managed.
What we recommend — and what we don't claim
No configuration can deliver this one, and we won't pretend otherwise. It is a human commitment: who sits on the review roster, whose judgement is sought before the AI is trusted with a decision, and whether the people affected can actually use the interface. What the deployment gives you is the raw material — the roster of named reviewers from Govern 2, and an audit trail showing whose judgement was actually applied, so the commitment can be checked rather than asserted.
What you can inspect
$ hermes audit --decisions --last 30d --group-by reviewer
j.doe    41 approvals   9 rejections
m.ruiz   12 approvals   1 rejection
# evidence of who reviews — the composition of that list is
# your organization's call, not the appliance's
Govern 4Requirement
Organizational teams are committed to a culture that considers and communicates AI risk — including surfacing problems rather than burying them.
What we recommend
Make the AI's own failures visible to the person in front of it, not just to a log nobody reads. The reality-check compares what the AI claimed it did against what actually happened and raises the warning in the session; flagging it files an incident in one step. A risk culture survives on how cheap it is to report something.
The configuration
# Post-turn verification, surfaced to the user — not silently logged
guardrails:
  reality_check: true
  surface_warnings: user      # user | log-only
feedback:
  flag_action: enabled          # one keystroke files an incident
  incidents: /var/hermes/incidents/
See it work
REALITY-CHECK WARNING  the assistant reported writing
  /matters/acme/summary.md — no such write occurred on disk.
[f] flag this   # → incidents/2026-07-30-0004.md, with the transcript
Govern 5Requirement · organizational practice
Processes are in place for robust engagement with the people affected by the AI system — users, staff, and outside parties.
What we recommend — and what we don't claim
Engagement is a conversation you hold, on a cadence you set; software cannot hold it for you. What we recommend is that the conversation be evidence-led rather than anecdotal: come to it with the incidents filed since the last review and the refusals the AI actually issued, so affected people are responding to what the system did, not to how it was described. The material is on the box already — the practice is putting it in front of the right people.
What you can inspect
$ hermes incidents --since 2026-04-01
2026-05-14  reality-check warning   resolved: prompt scope narrowed
2026-06-02  refused: out-of-scope    resolved: no action needed
2026-07-30  reality-check warning   open
# three items to walk through — the walking through is yours
Govern 6Requirement
Policies address the AI risks that arrive through third parties — the software, the data and the models you did not build yourself.
What we recommend
This is where a private deployment is structurally ahead: there is no third-party inference. The model runs locally, pinned to a specific digest, and egress is denied by default — so your prompts never become someone else's training data, and the supply chain is a list you can actually read to the end. Changing the model is a governed change, not a silent vendor update.
The configuration
# Local model, pinned by digest; no external inference path at all
model:
  source: local
  name: qwen3-35b
  digest: sha256:9f2c4b…      # pinned — a change needs approval
network:
  egress: deny                  # default-deny, not default-allow
  allow: []                   # no model API, no telemetry endpoint
See it work
$ hermes supply-chain
model ........ qwen3-35b  sha256:9f2c4b…  PINNED
external inference endpoints ............. 0
$ curl -s https://api.example-llm.com/v1  # from the appliance
curl: (7) egress denied by policy

Map

Knowing what you have actually deployed — its context, its category, its capabilities, and everything it touches. 5 categories

Map 1Requirement
The context is established and understood — intended purpose, setting, and who the system is for.
What we recommend
Write the purpose into the configuration and let it be enforceable: declare what the deployment is for, and have out-of-scope requests refused and logged rather than quietly attempted. A purpose that only exists in a slide deck cannot constrain anything.
The configuration
# Declared purpose, enforced as scope
deployment:
  purpose: "Draft and review internal legal documents"
  setting: "12-person firm, on-premises, no client PII export"
  in_scope:     [draft, summarize, search_internal]
  out_of_scope: [client_advice, external_publishing, hiring_decisions]
  on_out_of_scope: refuse_and_log
See it work
2026-07-30 11:04:12  identity=jdoe  action=draft
   target=hiring/candidate-rank.md  decision=REFUSE  policy=out-of-scope
# the refusal is recorded, so scope is auditable — not just intended
Map 2Requirement
The AI system is categorized — what task it performs, by what method, and what its knowledge limits are.
What we recommend
Keep a system card on the box, generated from the live configuration rather than typed by hand — task, model, method, knowledge cutoff, and an explicit list of what it cannot do. The limits matter more than the capabilities: most AI incidents start with someone assuming a capability that was never there.
The configuration
# Generated from live config — cannot drift from what is running
system_card:
  task: "retrieval-grounded drafting and summarization"
  method: "local LLM + local document index; no fine-tuning on client data"
  knowledge_cutoff: "2025-10"
  cannot: ["access the internet",
           "act without an approval for file changes",
           "see documents outside the granted role scope"]
See it work
$ hermes system-card
task ............ retrieval-grounded drafting and summarization
model ........... qwen3-35b (local)   cutoff 2025-10
cannot .......... internet · unapproved writes · out-of-role documents
generated from .. live config (policy v1.4)
Map 3Requirement
Capabilities, targeted usage, goals and expected benefits are understood — measured against the costs.
What we recommend
Express capability as an explicit allow-list of tools per role, with the benefit each one is there to deliver written next to it. Anything not on the list is unavailable — so "what can this AI do?" has a short, literal answer you can read off the box instead of inferring from a model's general abilities.
The configuration
# Capability = allow-list. Absent means unavailable, not merely discouraged.
tools:
  read_file:   {benefit: "find precedent fast",     roles: [legal-reader, ops-editor]}
  search_docs: {benefit: "answer from our own files", roles: [legal-reader, ops-editor]}
  write_file:  {benefit: "produce the first draft",  roles: [ops-editor], approval: required}
default: deny
See it work
$ hermes tools --role legal-reader
read_file      granted
search_docs    granted
write_file     not granted   # not "asks first" — unavailable
Map 4Requirement
Risks and benefits are mapped for all components of the AI system — including third-party software and data.
What we recommend
Reuse the Zero Trust work: because every capability is already its own Ziti service (800-207, tenet 1), the service list is the component inventory. Attach a risk note and an owner to each, then have the appliance cross-check the two lists — any service without a mapped risk is reported, so the inventory can't quietly fall behind the deployment.
The configuration
# Every reachable component carries a risk note and an owner
components:
  - service: private-ai
    risk: "generates text that may be wrong or over-confident"
    control: reality_check + approvals
    owner: "COO"
  - service: docs-index
    risk: "contains client-confidential matter files"
    control: role scope + Ziti dial policy
    owner: "Managing Partner"
  - service: tool-runner
    risk: "performs actions with side effects"
    control: approval gate (every)
    owner: "COO"
inventory_check: strict   # unmapped service = reported gap
See it work
$ hermes inventory --cross-check
ziti services ..... 3   mapped components ..... 3
unmapped .......... 0
# add a service without a risk note and this reads: unmapped 1 (gap)
Map 5Requirement · organizational practice
Impacts to individuals, groups, communities, organizations and society are characterized.
What we recommend — and what we don't claim
Judging who is affected, and how badly, is human work — it depends on your clients, your sector and your obligations, none of which the appliance knows. What we recommend is anchoring that judgement to two things the deployment does know: which classes of data the AI can reach, and which decisions it actually touched over a real period. That turns an impact assessment from a speculative exercise into a review of the record.
What you can inspect
$ hermes data-classes --reachable
client-confidential   reachable by: legal-reader, ops-editor
internal-general      reachable by: all roles
regulated-PII         reachable by: none   # excluded from the index
# who is affected, and how much it matters, is your assessment

Measure

Numbers you can actually compute on your own box — and evaluation that runs before a change ships, not after it breaks. 4 categories

Measure 1Requirement
Appropriate methods and metrics are identified and applied — and the choice of what you measure is itself documented.
What we recommend
Choose metrics that are computable from the audit log you already keep — refusal rate, how often the approval gate fires, how often the reality-check catches a false claim, how often answers are grounded in retrieved documents. No external telemetry, no vendor dashboard: the measurement stays on the appliance with the data.
The configuration
# Metrics computed locally from the audit log — nothing leaves the box
metrics:
  - {id: refusal_rate,       source: audit, window: 30d}
  - {id: approval_hit_rate,  source: audit, window: 30d}
  - {id: reality_check_rate, source: audit, window: 30d}
  - {id: grounded_rate,      source: audit, window: 30d}
report_to: /var/hermes/metrics/   # local file, not a cloud endpoint
See it work
$ hermes metrics --last 30d
refusal_rate ......... 3.1%   (41 of 1,320 requests)
approval_hit_rate .... 12.4%  (164 pauses for a human yes/no)
reality_check_rate ... 0.5%   (7 false claims caught)
grounded_rate ........ 94.2%  (answers citing a retrieved document)
Measure 2Requirement
The system is evaluated for trustworthy characteristics — and re-evaluated when it changes.
What we recommend
Keep a local evaluation set drawn from your own work and make it a gate, not a report: no model swap and no guardrail change ships unless the suite passes the thresholds you set. The point is that a regression is blocked before anyone relies on it, rather than discovered afterwards in an incident.
The configuration
# Evaluation gates the change — it does not merely describe it
eval:
  suite: /var/hermes/eval/firm-cases.yaml   # your documents, your tasks
  run_on: [model_change, guardrail_change, index_rebuild]
  thresholds:
    grounded_rate:   ">= 0.90"
    unsafe_refusal:  ">= 0.98"   # refuses what it should refuse
    regression_vs_last: "<= 0.02"
  on_fail: block_rollout
See it work
$ hermes eval run --candidate qwen3-35b-v2
grounded_rate ....... 0.87  threshold 0.90   FAIL
unsafe_refusal ...... 0.99  threshold 0.98   PASS
rollout ............. BLOCKED  # the old model stays in service
Measure 3Requirement
Mechanisms for tracking identified AI risks over time are in place — including risks that emerge after deployment.
What we recommend
Make the risk register a file, with each risk bound to the metric that watches it and a threshold that trips. A register nobody re-reads is the normal outcome; a register wired to a number means a drifting risk announces itself instead of waiting for the annual review.
The configuration
# Each risk names the metric that watches it and the level that trips
risks:
  - id: R1
    desc: "AI states something it did not do"
    watched_by: reality_check_rate
    trips_above: 0.01
    owner: COO
  - id: R2
    desc: "answers not grounded in our documents"
    watched_by: grounded_rate
    trips_below: 0.90
    owner: COO
  - id: R3
    desc: "scope creep into out-of-scope tasks"
    watched_by: refusal_rate
    trips_above: 0.10
    owner: "Managing Partner"
See it work
$ hermes risk status
R1  reality_check_rate  0.5%  trips >1.0%   OK    trend ↘
R2  grounded_rate      94.2%  trips <90%    OK    trend →
R3  refusal_rate        3.1%  trips >10%    OK    trend ↗ (watch)
Measure 4Requirement · organizational practice
Feedback about the efficacy of the measurement itself is gathered and assessed — are you measuring the right things?
What we recommend — and what we don't claim
A metric cannot audit itself. This category asks whether your numbers still describe reality — and only the people using the system can say. What we recommend is a short scheduled review that puts each metric next to the incidents of the same period and asks the one useful question: did anything go wrong that our numbers did not show? The appliance keeps the history locally so the comparison is possible; drawing the conclusion is yours.
What you can inspect
$ hermes metrics --history --by-quarter
                  Q1     Q2     Q3
reality_check    0.9%   0.7%   0.5%
incidents filed    4      3      3
# the metric improved while incidents held flat —
# exactly the discrepancy this review exists to notice

Manage

Acting on what you measured — prioritized responses, safe failure, governed third-party change, and a record of it all. 4 categories

Manage 1Requirement
Risks from mapping and measurement are prioritized, responded to, and managed — the highest first.
What we recommend
Bind each risk to a response that the configuration actually performs, and let severity pick the strength: high-severity risks get an enforced block, medium ones an approval pause, low ones an advisory warning. Priority that isn't expressed in the config is just a ranking in a spreadsheet.
The configuration
# Severity chooses the control strength — high risk gets a hard stop
responses:
  - {risk: R1, severity: high,   action: block}     # refuse to return the turn
  - {risk: R2, severity: medium, action: approve}   # pause for a human
  - {risk: R3, severity: low,    action: warn}      # advisory only
unhandled_risk: block   # a risk with no response fails closed
See it work
2026-07-30 15:41:02  identity=mruiz  action=write_file
   target=/matters/acme/summary.md  decision=BLOCK  policy=R1-high
# the claim failed the reality-check, so the turn was refused —
# not returned with a warning attached
Manage 2Requirement
Strategies maximize benefit and minimize negative impact — including fallback, deactivation and decommissioning plans.
What we recommend
Decide which way the system fails before it fails. We recommend fail-closed: if the guardrail config is missing, unreadable or unsigned, the AI refuses to serve rather than running ungoverned — the opposite of the common default. Pair it with a one-command stop and a written decommission path for the data.
The configuration
# Ungoverned operation is not a degraded mode — it is a stopped one
failure:
  mode: closed              # no guardrails => no service
  on_config_unreadable: refuse_all
kill_switch:
  command: "hermes stop --hard"   # drops Ziti bind + unloads model
decommission:
  index: wipe
  audit_log: "retain 7y"
  model: remove
See it work
$ mv /etc/hermes/guardrails.yaml /tmp && hermes doctor
guardrails ......... NOT FOUND
service ............ REFUSING ALL REQUESTS (fail-closed)
# it will not answer at all rather than answer ungoverned
Manage 3Requirement
Risks and benefits from third-party entities — models, software, data — are managed on an ongoing basis, not just at selection.
What we recommend
Treat the model like signed software: pinned by digest, unsigned artifacts refused, and any change routed through the approval gate and the evaluation suite. Because inference is local (Govern 6), a third party cannot change your system without you performing an act — which is precisely the ongoing management this category asks for.
The configuration
# A model change is a governed change — never a silent one
supply_chain:
  pin: strict                # digest must match the manifest
  allow_unsigned: false
  change_requires: [approval, eval_pass]
  manifest: /etc/hermes/manifest.lock
See it work
$ hermes model use ./mistral-new.gguf
digest sha256:41ab… not in manifest.lock ..... BLOCKED
required: approval by an owner + eval suite pass
# the running model is unchanged
Manage 4Requirement
Risk treatments are documented and monitored regularly, with plans for communicating incidents and responses.
What we recommend
Generate the monitoring report on the appliance, from the audit log — every treatment, whether it fired, and every incident with what changed afterwards. It stays local like everything else, and it is the document you hand an auditor or a client who asks how the AI is governed. Reporting that requires someone to assemble it by hand stops happening by about month four.
The configuration
# Monthly report generated on the box, from the record itself
reporting:
  cadence: monthly
  path: /var/hermes/reports/
  includes: [metrics, risk_status, incidents, approvals, config_changes]
  distribute_to: ["COO", "Managing Partner"]   # the comms plan
audit:
  enabled: true
  path: /var/hermes/audit.log        # stays on the appliance
See it work
$ hermes report --month 2026-07
risks tracked 3 · treatments fired 164 · incidents 3 (2 closed)
config changes 1 ... policy v1.3 → v1.4  approved by COO  2026-07-11
written to /var/hermes/reports/2026-07.md    local only

All nineteen categories of the NIST AI RMF Core — Govern, Map, Measure, Manage — each with the requirement, the approach we recommend, the configuration that implements it, and how you confirm it's working. Four are marked as organizational practice, where no configuration can honestly make the claim. This is standard 2 of the alignment set; standard 1 is NIST SP 800-207, Zero Trust.

Want this run against your setup?

The free Readiness Check tells you which of these you already meet and which need work — then the build puts the recommended configuration in place.

Start the free Readiness Check →