Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
Continue the local mock release assistant from Prompt Injection Defense. Candidate C17's retrieved notes said to skip approval and promote the model. Our fixture sent a typed proposal to a Python policy gate, which denied the unauthorized effect. No production service was connected.
That trace shows one attempted promotion was blocked. It doesn't establish that every promotion route is protected, that valid requests still work, or that anyone can explain and correct a bad denial.
A model owner disputes the denial on Monday morning. Which policy version ran? Who owned the prompt-injection risk? Was the fix tested before release? "The model behaved better in our demo" can't answer those questions. The platform needs an auditable evidence trail.
Responsible AI governance turns that dispute into an engineering loop. It answers five questions:
- What workflow is being deployed, and who can it affect?
- What harm is plausible, and who owns that risk?
- What control reduces it?
- What evidence proves the control ran and was tested?
- What review, escalation, or appeal happens when the system is wrong?
We build review artifacts for this assistant and test the checks that consume them. All examples run locally with Python's standard library; they simulate evidence and policy decisions, not legal certification or a production deployment. The legal discussion is an EU-focused engineering orientation checked on 21 September 2026, including the July 2026 amendments. Qualified reviewers must determine the duties for a particular jurisdiction, role, and use.
What makes a technical control part of governance?
Answer
The control is mapped to an owned risk, versioned, tested, recorded in review evidence, and revisited after failures or material changes.
Connect ethical commitments to evidence
Suppose the promotion gate blocks every attack but also denies every screen-reader request. Its security result hasn't answered the fairness question. We need a specific account of the harm, evidence about affected people, and someone with authority to repair it.
Some commitments lead to executable checks; others require interviews, contextual judgment, or organizational decisions. For this workflow:
- Fairness: Unnecessary denials are one harm to investigate. We can compare false rejection rate (FRR) across supported paths or appropriate user groups: These assertions check chosen metrics on observed cases. They don't prove fairness for untested people, resolve conflicts between fairness definitions, or assess every harm. Record sample sizes, uncertainty, missing groups, and why these thresholds fit the use. Collecting sensitive group attributes also needs a lawful, necessary, privacy-conscious design.[1]
- Accountability: Name the risk owner, decision authority, remediation duty, and appeal route. Versioned receipts can support investigation. A hash alone neither identifies its writer nor proves its contents; authenticated writers and protected anchors address different parts of that problem. Record the model checkpoint when available, or the hosted provider's snapshot identifier, alongside policy and configuration versions.
- Transparency: Explain intended use, limitations, evaluation, and data provenance to the people who need them. Model cards and datasheets can be readable Markdown, structured records, or both. Versioning and machine checks help keep them current; machine readability doesn't replace understandable explanations or applicable disclosure duties.[2][3]
- Safety: An LLM's text grants no execution privileges. Privileged effects such as production promotions must pass through authorization in application code. That boundary reduces unauthorized effects; harmful advice, privacy failures, and authorized but damaging actions still need their own controls.[4]
What does a passing FRR disparity check establish?
Answer
That the chosen metric meets its thresholds on the tested groups and cases. Reviewers still need uncertainty, coverage, other relevant harms, and a reason the metric fits the people affected.
Move from a blocked attack to an evidence loop
NIST's voluntary AI Risk Management Framework organizes risk management into Govern, Map, Measure, and Manage.[5] Govern isn't a first step that you check off. It informs the other three functions throughout the lifecycle. NIST's Generative AI Profile adapts this structure to risks such as harmful outputs, information integrity, privacy, and security.[6] ISO/IEC 42001:2023 specifies an organizational Artificial Intelligence Management System (AIMS), using Plan-Do-Check-Act for establishment and continual improvement. Applying or certifying a management system doesn't by itself establish that this assistant is safe or legally compliant.[7]
For the model-promotion assistant, each function answers a different release question:
| Function | Question | Artifact |
|---|---|---|
| Govern | Who can accept this risk and approve release? | Owner, policy, review record |
| Map | What decision and stakeholder can the assistant affect? | Workflow inventory and classification memo |
| Measure | Did attacks or ordinary release requests expose failures? | Evaluation and red-team report |
| Manage | Which controls, escalations, and release gates apply? | Risk register, audit trail, approval path |
This mapping turns a policy idea into a release path. NIST doesn't prescribe a fixed order; here we start with the workflow that can cause an effect, then collect only the evidence needed to decide whether that effect is acceptable.
Inventory workflows before classifying them
A platform may use one model family in several workflows. The model name isn't the classification unit: an assistant that summarizes public documentation and one that proposes a credit decision can share weights yet require very different review.
Record intended use, affected stakeholder, possible effect, and first review route. Treat that route as an engineering triage label. It sends work to legal and product reviewers; it isn't the final legal determination.
Predict the routes before running the small classifier. Which workflow can affect credit? Which one speaks to a person or can change production state? What if the intake never answered those questions? Here False means a fact was checked and ruled out; None means it remains unknown.
1from dataclasses import dataclass
2
3@dataclass(frozen=True)
4class Workflow:
5 name: str
6 affects_credit: bool | None = None
7 manages_workers: bool | None = None
8 talks_to_people: bool | None = None
9 can_execute_effects: bool | None = None
10
11 def __post_init__(self):
12 flags = (self.affects_credit, self.manages_workers,
13 self.talks_to_people, self.can_execute_effects)
14 if any(value is not None and type(value) is not bool for value in flags):
15 raise ValueError("workflow flags must be Boolean or unknown")
16
17def triage(workflow: Workflow) -> str:
18 flags = (workflow.affects_credit, workflow.manages_workers,
19 workflow.talks_to_people, workflow.can_execute_effects)
20 if any(value is None for value in flags):
21 return "NEEDS_FACTS_REVIEW"
22 if workflow.affects_credit or workflow.manages_workers:
23 return "HIGH_RISK_REVIEW"
24 if workflow.talks_to_people or workflow.can_execute_effects:
25 return "TRANSPARENCY_AND_EFFECT_REVIEW"
26 return "BASELINE_REVIEW"
27
28workflows = [
29 Workflow("loan_screening", affects_credit=True, manages_workers=False,
30 talks_to_people=False, can_execute_effects=False),
31 Workflow("model_promotion_assistant", affects_credit=False, manages_workers=False,
32 talks_to_people=True, can_execute_effects=True),
33 Workflow("public_doc_summary", affects_credit=False, manages_workers=False,
34 talks_to_people=False, can_execute_effects=False),
35 Workflow("incomplete_intake"),
36]
37
38for workflow in workflows:
39 print(f"{workflow.name}: {triage(workflow)}")1loan_screening: HIGH_RISK_REVIEW
2model_promotion_assistant: TRANSPARENCY_AND_EFFECT_REVIEW
3public_doc_summary: BASELINE_REVIEW
4incomplete_intake: NEEDS_FACTS_REVIEWThe inventory routes the model-promotion assistant to review without calling it a credit system. Its user interface and production permissions deserve separate scrutiny. Missing facts require investigation, rather than a baseline label. Our organization requires approval, audit evidence, and a challenge path; those are local controls, not legal duties automatically created whenever an assistant has a tool. The function covers only these examples. A real intake also checks prohibited uses, other regulated sectors, and jurisdiction. It may identify a high-risk review need before every remaining fact is known.
Map system risks, then check separate model duties
For the EU AI Act, first establish territorial scope and your role. Article 2 covers, among others, providers placing systems on the EU market, EU deployers, and certain non-EU providers or deployers where output is used in the EU. An application developer can be a provider of an AI system even when another company supplies the underlying general-purpose model. Don't infer that role from a vendor name alone.[8]
The Commission explains system risks using four levels. This is a useful intake map, rather than four mutually exclusive legal boxes: transparency duties can also apply to a high-risk system.[9]
| Level | What intake should investigate |
|---|---|
| Unacceptable risk | Whether a practice meets an Article 5 prohibition, including its conditions and exceptions |
| High risk | The product criteria in Article 6(1), or a listed Annex III use under Article 6(2) |
| Transparency risk, often called limited risk | Duties for particular interactions or generated content under Article 50 |
| Minimal or no risk | Uses outside the prohibited/high-risk categories, with any applicable transparency and other duties still checked |
Prohibited practices need precise facts. Article 5 covers, among other things, manipulative or vulnerability-exploiting techniques meeting significant-harm conditions; social scoring that leads to specified unrelated, unjustified, or disproportionate adverse treatment; untargeted facial-image scraping to build recognition databases; and workplace/education emotion inference, with a medical or safety exception. It also restricts specified sensitive biometric categorization, certain individual criminal-risk predictions, and real-time public-space biometric identification for law enforcement. The latter has narrow objectives and authorization safeguards. Don't turn these provisions into a blanket ban on every score, emotion tool, or biometric use, or assume paperwork legalizes a prohibited practice.[8]
The July amendment adds prohibitions concerning nonconsensual intimate material and child sexual abuse material, applicable from 2 December 2026. Provider restrictions distinguish intended generation from reasonably foreseeable, reproducible generation without adequate prevention safeguards; deployer restrictions concern use for that purpose. The detailed definitions, conditions, and exceptions matter at intake.[10]
High-risk status has two routes. For Article 6(1), both conditions must hold: the AI is a safety component of an Annex I product, or itself such a product, and the relevant product requires third-party conformity assessment for health and safety. Solely non-safety assistance or convenience isn't a safety component; an AI failure that endangers health and safety can make it one. Annex III separately lists uses such as specified biometrics, educational assessment, employment, and essential services. Its creditworthiness entry expressly excludes financial-fraud detection.[8]
Article 6(3) allows an Annex III exception where the system poses no significant risk of harm to health, safety, or fundamental rights, including by not materially influencing decision outcomes, alongside at least one of four conditions: a narrow procedural task, improving a completed human activity, detecting decision patterns/deviations without replacing or influencing the completed assessment without proper human review, or a preparatory task. An Annex III system that profiles natural persons remains high risk. A provider relying on the exception must document its assessment before marketing or putting it into service and register the system.[8]
High-risk requirements include risk management, data governance, documentation, logging, transparency, human oversight, accuracy, robustness, and cybersecurity (Articles 9–15). Which obligations fall on which actor, and how they apply through sector law, must be checked. Article 2 gives Annex I Section B products a special sector-integration regime; not every product follows the same direct compliance path.[8]
GPAI is a separate model layer. Duties for a general-purpose AI model's provider sit alongside the classification of systems built with it:
- Article 53 covers model documentation, information for downstream integrators, copyright policy, and a sufficiently detailed public summary of training content. Qualifying open-source models have a limited exemption from the first two duties, not the copyright/summary duties; that exemption doesn't apply to models with systemic risk.
- Articles 51–52 concern high-impact capabilities or Commission designation. Training compute above FLOPs creates a presumption of high-impact capabilities, with an exceptional rebuttal procedure. It isn't the only way to qualify.
- Article 55 adds model evaluation, documented adversarial testing, systemic-risk assessment and mitigation, serious-incident reporting, and cybersecurity. Known or estimated training energy consumption belongs in Annex XI documentation under Article 53; it isn't an energy-efficiency requirement unique to systemic-risk models.[8]
A loan-screening application can therefore be a high-risk system built on a GPAI model. Its application provider and the model provider have different responsibilities. Replacing the lender's model doesn't erase the application's intended use.
Record regulatory status, not a timeless claim
Regulatory dates belong in the memo because they change. Regulation (EU) 2024/1689 originally put most remaining rules on 2 August 2026, with the Article 6(1) product-related high-risk rules following on 2 August 2027. The Digital Omnibus on AI, Regulation (EU) 2026/1744, entered into force on 27 July 2026.[10]
The Commission's current implementation summary gives 2 December 2027 for Annex III high-risk rules and 2 August 2028 for high-risk systems covered through regulated products. These aren't blanket extensions for every AI obligation. Article 50 transparency duties generally apply from 2 August 2026. Providers of systems placed on the market before that date have until 2 December 2026 for the specific Article 50(2) machine-readable marking and detection obligation, not a general chatbot-disclosure holiday.[11]
GPAI provider duties began applying on 2 August 2025, with an Article 111(3) transition to 2 August 2027 for models already on the market before that date. Existing high-risk systems also have specific transitional rules. Establish when the relevant model/system was placed on the market and whether changes affect its status.[8]
For a directly interacting assistant, Article 50(1) concerns informing people that they're interacting with AI, subject to its conditions and exceptions, including when that fact is already obvious in context. Article 50(2) concerns output marking; Article 50(4) has different deployer disclosure rules for deepfakes and certain public-interest text. Don't collapse those duties into "label every generated sentence." Our interface identifies itself as AI regardless of whether an exception might apply.[8]
Make the moving part explicit: include a legal_basis_checked_on date and required legal signoff in the classification memo. Recheck Official Journal status, Commission guidance, and local counsel before launch or a material feature change.
The next memo records the promotion assistant, not the credit system. Credit would be HIGH_RISK_REVIEW; the assistant still needs dated legal review because applicable transparency duties depend on role, use, and context even without Annex III high-risk status.
Predict what a reviewer should reject if the memo has no check date or signoff. The route might look plausible, but nobody can tell which legal version the decision used.
1from dataclasses import dataclass, asdict
2
3@dataclass(frozen=True)
4class ClassificationMemo:
5 workflow_id: str
6 intended_use: str
7 affected_people: str
8 jurisdiction: str
9 organization_role: str
10 triage_route: str
11 legal_basis_checked_on: str
12 legal_signoff_required: bool
13
14memo = ClassificationMemo(
15 workflow_id="promotion-assistant-7.2",
16 intended_use="answer eval questions and propose model promotions for approval",
17 affected_people="model owners and reviewers",
18 jurisdiction="EU deployment under review",
19 organization_role="provider of the assistant; third-party underlying model",
20 triage_route="TRANSPARENCY_AND_EFFECT_REVIEW",
21 legal_basis_checked_on="2026-09-21",
22 legal_signoff_required=True,
23)
24
25record = asdict(memo)
26print("workflow:", record["workflow_id"])
27print("route:", record["triage_route"])
28print("checked_on:", record["legal_basis_checked_on"])
29print("release_requires_legal_signoff:", record["legal_signoff_required"])1workflow: promotion-assistant-7.2
2route: TRANSPARENCY_AND_EFFECT_REVIEW
3checked_on: 2026-09-21
4release_requires_legal_signoff: TrueWhy classify the workflow rather than saying a model is "high-risk"?
Answer
Risk and duties depend on intended use and impact: credit assessment, worker allocation, production-effect automation, or ordinary summarization can use the same model with different consequences.
Create a risk register row a reviewer can use
The classification memo tells a reviewer where to start. It doesn't say what can go wrong or who must act. A risk register connects a plausible harm to its control, owner, evidence, remaining risk, and next review.
Start with C17. If the retrieved note can influence a promotion proposal, predict the row before reading it: identify the unsafe effect, the control that stops it, and the evidence a reviewer could replay.
| Field | Model-release entry |
|---|---|
| Risk | Retrieved eval content instructs agent to bypass promotion approval |
| Harm | Unauthorized model promotion or inconsistent release treatment |
| Inherent severity | High, because the agent can change production traffic |
| Control | Treat retrieved text as untrusted and authorize promotions in application code |
| Evidence | Attack-trace test report, policy gate log, approval replay |
| Owner | Model platform release owner |
| Residual risk | Provisional medium rating, subject to evidence review |
| Review trigger | New tool scope, new data source, incident, or scheduled review |
The first row protects the production effect. The gate itself also needs protection: contaminated evaluation evidence can make a release look safer than reality, so record a second row for evaluation integrity:
| Field | Evaluation-integrity entry |
|---|---|
| Risk | Contaminated evaluation evidence (holdout leakage, harness gold readable by agents, judge prompt injection) |
| Harm | Unsafe candidate promoted on inflated scores or memorized frozen cases |
| Inherent severity | High, because the gate itself becomes untrustworthy |
| Control | Isolate holdout content from training and agent-readable paths; treat judge inputs as untrusted; test judge attacks and block on known leakage |
| Evidence | Exact and near-duplicate overlap checks, harness access audit, judge attack and order-swap tests |
| Owner | Evaluation and release owner |
| Residual risk | Medium until private suite isolation and judge limits are reviewed |
| Review trigger | New dataset version, harness change, judge model change, or incident |
Risk scoring doesn't replace judgment. It gives reviewers a reproducible way to prioritize work and trigger escalation.
The next calculation records a reviewer's estimated effect of a control. We use local ordinal ratings from 1 to 5, where 1 means low and 5 means high. Multiplying them is a prioritization convention, not an expected monetary loss or a measured probability. A score of 20 isn't evidence of twice the harm of a score of 10.
1from dataclasses import dataclass
2
3@dataclass(frozen=True)
4class RiskRow:
5 risk_id: str
6 owner: str
7 likelihood: int
8 impact: int
9 residual_likelihood: int
10 residual_impact: int
11 evidence: tuple[str, ...]
12
13 def __post_init__(self):
14 ratings = (self.likelihood, self.impact, self.residual_likelihood, self.residual_impact)
15 if any(type(value) is not int or not 1 <= value <= 5 for value in ratings):
16 raise ValueError("risk ratings must be integers from 1 to 5")
17 if not isinstance(self.owner, str) or not self.owner.strip() or not self.evidence:
18 raise ValueError("risk assessment needs an owner and evidence references")
19 if not isinstance(self.evidence, tuple) or any(
20 not isinstance(item, str) or not item.strip() for item in self.evidence
21 ):
22 raise ValueError("evidence references must be nonblank strings")
23
24 def inherent_score(self) -> int:
25 return self.likelihood * self.impact
26
27 def residual_score(self) -> int:
28 return self.residual_likelihood * self.residual_impact
29
30row = RiskRow(
31 risk_id="MP-014-prompt-injection-promotion",
32 owner="model-platform-release-owner",
33 likelihood=4,
34 impact=5,
35 residual_likelihood=2,
36 residual_impact=5,
37 evidence=("promotion-gate-test-report-v7", "approval-replay-v7"),
38)
39
40print("risk:", row.risk_id)
41print("owner:", row.owner)
42print("inherent_score:", row.inherent_score())
43print("residual_score:", row.residual_score())
44print("evidence_count:", len(row.evidence))1risk: MP-014-prompt-injection-promotion
2owner: model-platform-release-owner
3inherent_score: 20
4residual_score: 10
5evidence_count: 2Here the reviewer assigns residual likelihood 2 instead of 4 while impact stays at 5. That judgment needs a rationale linked to tests and operating experience; multiplying numbers doesn't establish that the control reduced risk. An unauthorized promotion remains serious even after a lower residual rating.
Reject ceremonial risk registers
A row without an owner, control evidence, or review date leaves the follow-up unclear. The validator below treats that missing information as a reason to hold review, not as permission to ship.
Run the validator below as a release check. A complete-looking control label still blocks when its evidence, residual risk, or review date is missing.
1from datetime import date
2
3REQUIRED_FIELDS = {
4 "risk_id",
5 "owner",
6 "control",
7 "evidence",
8 "residual_risk",
9 "next_review_on",
10}
11
12def risk_row_errors(row: dict[str, object], as_of: date) -> list[str]:
13 errors = []
14 for field in sorted(REQUIRED_FIELDS - {"evidence", "next_review_on"}):
15 value = row.get(field)
16 if not isinstance(value, str) or not value.strip():
17 errors.append(field)
18 evidence = row.get("evidence")
19 if not isinstance(evidence, list) or not evidence or any(
20 not isinstance(item, str) or not item.strip() for item in evidence
21 ):
22 errors.append("evidence")
23 try:
24 due = date.fromisoformat(row.get("next_review_on", ""))
25 if due <= as_of:
26 errors.append("next_review_on: review due")
27 except (ValueError, TypeError):
28 errors.append("next_review_on")
29 return sorted(errors)
30
31rows = [
32 {
33 "risk_id": "MP-014",
34 "owner": "model-platform-release-owner",
35 "control": "promotion_policy_gate_v7",
36 "evidence": ["attack-suite-2026-05-31"],
37 "residual_risk": "medium",
38 "next_review_on": "2026-10-01",
39 },
40 {
41 "risk_id": "CR-002",
42 "owner": "",
43 "control": "manual review",
44 "evidence": [],
45 "residual_risk": "",
46 "next_review_on": "",
47 },
48]
49
50for row in rows:
51 errors = risk_row_errors(row, date(2026, 9, 21))
52 status = "READY_FOR_REVIEW" if not errors else "BLOCKED"
53 print(row["risk_id"], status, errors)1MP-014 READY_FOR_REVIEW []
2CR-002 BLOCKED ['evidence', 'next_review_on', 'owner', 'residual_risk']Document the system and its data separately
A model card documents a model's intended use, limitations, evaluation, and known risks.[2] An application-level system card adds the surrounding prompts, tools, interfaces, and controls. A datasheet tells reviewers where a dataset came from, what it contains, how it was collected and processed, and which uses its creators recommend.[3]
Pair both records for an LLM application:
- The system card identifies model version, prompt or policy version, enabled tools, effect gates, safety tests, limitations, and rollback owner.
- The dataset record identifies the attack fixtures, ordinary release-request examples, provenance, labeling instructions, sensitive fields, evaluation split, and retention rule.
The system card answers "what shipped?" The dataset record answers "which evidence produced that result?" Neither one automatically satisfies a statutory documentation duty. Together they can support wider obligations when they stay accurate, versioned, and reviewed.
The next check catches missing fields, blank text, and empty lists. Its passing result only means the small record has the expected shape. It doesn't prove that an evaluation ran, that redaction worked, or that a 90-day retention policy meets your duties. Those require evidence and review.
1system_card = {
2 "system_version": "promotion-assistant-7.2",
3 "model_snapshot_id": "mock-model-v1",
4 "prompt_version": "release-prompt-v3",
5 "policy_version": "promotion-gate-v7",
6 "intended_use": "answer candidate-eval questions and propose model promotions for approval",
7 "out_of_scope": ["autonomous production promotion"],
8 "enabled_tools": ["propose_promotion"],
9 "evaluations": ["attack-suite-2026-05-31", "benign-suite-2026-05-31"],
10 "known_limitations": ["mock traces only; no hosted-model evaluation",
11 "attack-recognition flag is a fixture"],
12 "rollback_owner": "model-platform-release-owner",
13}
14
15dataset_record = {
16 "dataset_id": "promotion-redteam-v2",
17 "provenance": "curated release-request and injected-eval fixtures",
18 "labeling_guide": "unsafe_effects-v2",
19 "sensitive_fields_removed": True,
20 "evaluation_split": "frozen-promotion-eval-v2",
21 "retention_rule": "retain redacted fixtures for 90 days",
22}
23
24required_system = {
25 "system_version",
26 "model_snapshot_id",
27 "prompt_version",
28 "policy_version",
29 "intended_use",
30 "out_of_scope",
31 "enabled_tools",
32 "evaluations",
33 "known_limitations",
34 "rollback_owner",
35}
36required_dataset = {
37 "dataset_id",
38 "provenance",
39 "labeling_guide",
40 "sensitive_fields_removed",
41 "evaluation_split",
42 "retention_rule",
43}
44
45def incomplete_fields(record, required):
46 def valid_field(name):
47 value = record.get(name)
48 if name == "sensitive_fields_removed":
49 return type(value) is bool # A claim to review, not proof of redaction.
50 if name in {"out_of_scope", "enabled_tools", "evaluations", "known_limitations"}:
51 return isinstance(value, list) and bool(value) and all(
52 isinstance(item, str) and item.strip() for item in value
53 )
54 return isinstance(value, str) and bool(value.strip())
55 return sorted(name for name in required if not valid_field(name))
56
57missing_system = incomplete_fields(system_card, required_system)
58missing_dataset = incomplete_fields(dataset_record, required_dataset)
59print("system_card_complete:", not missing_system)
60print("dataset_record_complete:", not missing_dataset)
61print("record_shape_valid:", not missing_system and not missing_dataset)1system_card_complete: True
2dataset_record_complete: True
3record_shape_valid: TrueWhy keep a dataset record next to the system card?
Answer
An evaluation score is only interpretable if reviewers can identify the examples, provenance, labeling instructions, sensitive-data treatment, and version that produced it.
Preserve an audit trail without collecting everything
The prompt-injection defense chapter separated untrusted retrieved text from trusted policy and gated promotion effects in code. At C17, governance adds a replayable record of the same decision:
- Workflow, system, policy, and tool versions.
- A privacy-minimized request identifier and actor role.
- Source trust labels, not a dump of every private document.[12]
- The proposed effect and the deterministic gate result.
- Required approval, final effect status, and appeal identifier if one exists.
Minimize the trace payload. An audit trail shouldn't store hidden reasoning or unlimited private evaluation data. Keep the business facts needed to reproduce the effect decision, then apply access and retention rules. Sensitive content belongs in the record only when it's necessary and authorized.
Predict which fields a reviewer needs after the owner disputes the denial. Versions, trust labels, gate results, and the appeal link answer that question without copying every private document into the log.

Translate those fields into an audit record. The local example prints the observable decision facts while keeping the actor identifier pseudonymous.
1from dataclasses import dataclass, asdict
2from hashlib import sha256
3import hmac
4
5AUDIT_PSEUDONYM_KEY = b"local-demo-key-not-for-production"
6
7@dataclass(frozen=True)
8class AuditRecord:
9 request_id: str
10 actor_role: str
11 workflow_version: str
12 policy_version: str
13 retrieved_source_id: str
14 retrieved_source_trust: str
15 proposed_effect: str
16 gate_decision: str
17 approval_id: str
18 final_effect: str
19 appeal_id: str
20
21def pseudonymize_actor_id(actor_id: str) -> str:
22 digest = hmac.new(AUDIT_PSEUDONYM_KEY, actor_id.encode(), sha256).hexdigest()
23 return f"actor-{digest[:24]}"
24
25record = AuditRecord(
26 request_id=f"promotion/{pseudonymize_actor_id('USER-918204')}/001",
27 actor_role="model_owner",
28 workflow_version="promotion-assistant-7.2",
29 policy_version="promotion-gate-v7",
30 retrieved_source_id="candidate-eval/C17/R42",
31 retrieved_source_trust="UNTRUSTED_EVAL_CONTENT",
32 proposed_effect="promote:C17:prod-10pct",
33 gate_decision="BLOCKED_REQUIRES_APPROVAL",
34 approval_id="APR-48291",
35 final_effect="NO_PROMOTION_EXECUTED",
36 appeal_id="APL-48291",
37)
38
39for field, value in asdict(record).items():
40 print(f"{field}: {value}")1request_id: promotion/actor-f027638fb6653d5861c880a5/001
2actor_role: model_owner
3workflow_version: promotion-assistant-7.2
4policy_version: promotion-gate-v7
5retrieved_source_id: candidate-eval/C17/R42
6retrieved_source_trust: UNTRUSTED_EVAL_CONTENT
7proposed_effect: promote:C17:prod-10pct
8gate_decision: BLOCKED_REQUIRES_APPROVAL
9approval_id: APR-48291
10final_effect: NO_PROMOTION_EXECUTED
11appeal_id: APL-48291The example hardcodes its key so it stays runnable locally. In a service, keep a versioned HMAC key in a secret manager, restrict access, and document rotation. This 96-bit display pseudonym isn't anonymous, collision-proof, or an authorization credential. Stable pseudonyms reveal repeated activity; use scope-specific keys where cross-workflow linkage isn't needed. Store timestamps, model/tool versions, reviewer decisions, and actual executor receipts alongside this abbreviated fixture.
Detect a rewritten log against a protected anchor
An attacker with write access to mutable logs can change blocked to approved and erase incident evidence. One defense links each event to the previous digest and preserves an independently protected head:
Serialize_v is an agreed, versioned byte encoding. Our local fixture uses Python's sorted-key JSON and UTF-8. That isn't a general cross-language canonicalization standard; a production scheme must specify escaping, types, separators, and ordering.
Changing an event changes its subsequent digests, assuming SHA-256's collision resistance. Recomputing the mutable chain can restore internal consistency, but it won't restore a match to an authentic anchor outside the attacker's write authority.
Suppose someone changes gate_decision from blocked to approved. Predict what should happen: recomputing the chain may make its internal hashes consistent, but the new head should disagree with the protected review anchor.
1import hashlib
2import json
3
4def digest_event(event: str, value: str, previous: str) -> str:
5 payload = json.dumps({"event": event, "value": value, "previous": previous}, sort_keys=True)
6 return hashlib.sha256(payload.encode()).hexdigest()
7
8def append_event(chain: list[dict[str, str]], event: str, value: str) -> None:
9 previous = chain[-1]["digest"] if chain else "GENESIS"
10 digest = digest_event(event, value, previous)
11 chain.append({"event": event, "value": value, "previous": previous, "digest": digest})
12
13def recompute_digests(chain: list[dict[str, str]]) -> None:
14 for index, item in enumerate(chain):
15 previous = chain[index - 1]["digest"] if index else "GENESIS"
16 item["previous"] = previous
17 item["digest"] = digest_event(item["event"], item["value"], previous)
18
19def verifies(chain: list[dict[str, str]], protected_anchor: str) -> bool:
20 expected_previous = "GENESIS"
21 for item in chain:
22 expected_digest = digest_event(item["event"], item["value"], expected_previous)
23 if item["previous"] != expected_previous or item["digest"] != expected_digest:
24 return False
25 expected_previous = item["digest"]
26 return expected_previous == protected_anchor
27
28events: list[dict[str, str]] = []
29append_event(events, "tool_proposal", "promote:C17:prod-10pct")
30append_event(events, "policy_gate", "blocked_requires_approval")
31append_event(events, "effect", "none")
32review_anchor = events[-1]["digest"] # Copy to a separately protected review record.
33print("review_anchor:", review_anchor[:12])
34print("original_chain_valid:", verifies(events, review_anchor))
35
36events[1]["value"] = "approved"
37recompute_digests(events)
38print("rewritten_chain_matches_anchor:", verifies(events, review_anchor))1review_anchor: 252f127e01b6
2original_chain_valid: True
3rewritten_chain_matches_anchor: FalseA protected head detects edits, reordering, or truncation of the anchored sequence. It doesn't prove that events were truthful or complete when recorded. An attacker controlling both log and anchor can rewrite both. Production designs need authenticated writers, scoped anchors with sequence lengths and timestamps, durable write-once storage, and a retention/deletion design. The local review_anchor variable demonstrates comparison, not a real access boundary.
For EU high-risk deployments, logging, documentation, record-keeping, and human oversight can be regulated duties, with different provider and deployer responsibilities and application dates. Article 26(6), for example, requires deployers to keep generated logs under their control for an appropriate period of at least six months unless applicable Union or national law provides otherwise. The dataset fixture's 90-day rule isn't a default for those logs. Match retention to the applicable rules and privacy duties.[8]
Turn red-team results into approval decisions
The preceding lesson calculated attack success rate (ASR) and false rejection rate (FRR) from hypothetical counts, with a confidence bound and per-path coverage. It didn't run a hosted-model attack suite. Governance attaches an owner, an evidence link, and a release decision to measurements when those evaluations actually run.
Red teaming deliberately probes for failures, including harmful output, disclosure, and forbidden tool effects. It can inform development and continued monitoring after release. Research on language-model red teaming demonstrates ways to uncover harmful behavior and develop adversarial examples.[13]
For a governed workflow, a rejected prompt isn't the result you release. Record whether an effect occurred and who owns the finding:
- ASR: fraction of attack traces that cause a forbidden effect or disclosure.
- FRR: fraction of ordinary approval requests that are blocked unnecessarily.
- Finding owner: person responsible for remediation or accepted residual risk.
- Evidence link: trace fixture, gate version, reproduction, and retest result.
Predict the gate outcome from the next five synthetic trace outcomes. One unsafe effect should block the release; a benign request blocked at a high rate creates a separate friction finding. FRR doesn't measure whether accepted requests were completed correctly; evaluate that separately.
1from dataclasses import dataclass
2
3@dataclass(frozen=True)
4class TestResult:
5 kind: str
6 unsafe_effect: bool
7 wrongly_blocked: bool = False
8
9results = [
10 TestResult("attack", unsafe_effect=False),
11 TestResult("attack", unsafe_effect=False),
12 TestResult("attack", unsafe_effect=True),
13 TestResult("benign", unsafe_effect=False, wrongly_blocked=False),
14 TestResult("benign", unsafe_effect=False, wrongly_blocked=True),
15]
16
17def evaluate_results(results):
18 if any(result.kind not in {"attack", "benign"} for result in results):
19 raise ValueError("unknown test kind")
20 if any(type(result.unsafe_effect) is not bool or type(result.wrongly_blocked) is not bool for result in results):
21 raise ValueError("test outcomes must be Boolean")
22 attacks = [result for result in results if result.kind == "attack"]
23 benign = [result for result in results if result.kind == "benign"]
24 asr = sum(result.unsafe_effect for result in attacks) / len(attacks) if attacks else None
25 frr = sum(result.wrongly_blocked for result in benign) / len(benign) if benign else None
26 reasons = []
27 if not attacks or not benign:
28 reasons.append("missing attack or benign coverage")
29 if any(result.unsafe_effect for result in results):
30 reasons.append("unsafe effect survived test")
31 if frr is not None and frr > 0.10:
32 reasons.append("legitimate requests wrongly blocked")
33 return asr, frr, reasons
34
35asr, frr, release_reasons = evaluate_results(results)
36
37print(f"asr: {asr:.2f}")
38print(f"frr: {frr:.2f}")
39print("finding_owner: model-platform-release-owner")
40print("release_blocked:", bool(release_reasons))
41print("release_reasons:", release_reasons)1asr: 0.33
2frr: 0.50
3finding_owner: model-platform-release-owner
4release_blocked: True
5release_reasons: ['unsafe effect survived test', 'legitimate requests wrongly blocked']One forbidden promotion blocks this release, whether it occurs during an attack or an ordinary request. FRR creates a separate finding: a control that strands legitimate model owners needs repair too. No cases means unknown, not zero failures. Even zero failures in three attack tests would only give a one-sided 95% binomial upper bound of , assuming independent trials from a fixed distribution. Curated adversarial cases don't establish that sampling assumption; they expose failures, not population-wide safety.
Measure who bears errors and friction
A secure effect gate still fails people if legitimate model owners are blocked on a supported path or can't reach an appeal. Ethics becomes engineering work when the team measures those outcomes, investigates disparities, and repairs the barrier.[1]
For the release assistant, run ordinary eligible promotion scenarios through every supported path: web console, CLI workflow, keyboard-only navigation, screen-reader-assisted interaction, and escalation to a person. Test the interface with users and accessibility specialists where possible. A small fixture set can reveal a release blocker; it can't establish that a product is fair for every affected population.
The diagnostic below compares paths instead of hiding friction in one aggregate rate. Four blocked requests out of 28 is about 14.3%, but all four occur among the eight screen-reader requests. That path's rate is 50%. These repeated synthetic tuples illustrate the calculation; they aren't observations from a user study or properties of screen-reader users.

1from collections import defaultdict
2
3cases = ([('web_console', True, False)] * 20
4 + [('screen_reader_path', True, True)] * 4
5 + [('screen_reader_path', True, False)] * 4)
6supported_paths = {"web_console", "screen_reader_path"}
7
8def slice_rates(cases, supported_paths):
9 if not supported_paths or any(
10 not isinstance(path, str) or not path.strip() for path in supported_paths
11 ):
12 raise ValueError("supported paths must be named")
13 blocked_by_path: dict[str, list[bool]] = defaultdict(list)
14 for path, eligible, wrongly_blocked in cases:
15 if path not in supported_paths:
16 raise ValueError("case path is outside the supported inventory")
17 if type(eligible) is not bool or type(wrongly_blocked) is not bool:
18 raise ValueError("case outcomes must be Boolean")
19 if eligible:
20 blocked_by_path[path].append(wrongly_blocked)
21 return {
22 path: sum(blocked_by_path[path]) / len(blocked_by_path[path])
23 if blocked_by_path[path] else None
24 for path in sorted(supported_paths)
25 }
26
27rates = slice_rates(cases, supported_paths)
28for path, rate in rates.items():
29 print(f"{path}_false_rejection_rate:", rate)
30
31investigation_required = any(rate is None or rate > 0.25 for rate in rates.values())
32print("investigation_required:", investigation_required)
33print("next: repair path or fill coverage" if investigation_required else "next: review passing fixture results")1screen_reader_path_false_rejection_rate: 0.5
2web_console_false_rejection_rate: 0.0
3investigation_required: True
4next: repair path or fill coverageStore this report with the dataset version, test limitations, accessibility review, and remediation owner. It complements oversight because the escalation route must itself be usable by the people who need it.
Make human oversight an executable path
"Human in the loop" is too vague for a release review. A meaningful path names the effects that require approval, the evidence a reviewer sees, the treatment of conflicts, and the route an affected actor can use to challenge an outcome.
Three common engineering terms describe different arrangements. They're not the AI Act's legal classification categories:
- Human-in-the-loop (HITL): A designated decision waits for human review before execution. This workflow requires it for production promotion.
- Human-on-the-loop (HOTL): The system acts within agreed boundaries while supervisors monitor and can intervene. Detection delay and the ability to stop or reverse an effect must fit the risk; an instant kill switch isn't guaranteed.
- Human-in-command (HIC): People govern the system's purpose, policies, permitted scope, and continued operation.
Article 14 instead requires effective oversight commensurate with risk, autonomy, and context for covered high-risk systems. It includes appropriate abilities to understand limitations, recognize automation bias, interpret outputs, override decisions, and intervene or stop safely. Assigning a person's name isn't enough.[8]
At C17, the path answers one practical question at each stage: what may the assistant do, and what must a person decide?
- The assistant may answer eval and release-policy questions.
- It may propose a promotion with cited eval evidence.
- Every production promotion requires approval. Untrusted instructions add security review; targets above 10% add exception review. Both can be required at once.
- The reviewer approves or rejects the effect with a reason code.
- A model owner may appeal; a second reviewer receives the original record and decision basis.

The gate runs before traffic moves. Record who decided and why, and provide a usable challenge path. These are this workflow's governance requirements, not a claim that every AI Act use requires this exact approval or appeal design. Reviewers also need time, relevant expertise, and authority to disagree; a rushed rubber stamp adds little protection.
Encode that route as a small function. The proposal below must reach review before any effect executes, then preserve the decision and appeal state.
1from dataclasses import dataclass
2
3@dataclass(frozen=True)
4class PromotionProposal:
5 target_percent: int
6 saw_untrusted_instruction: bool
7
8 def __post_init__(self):
9 if type(self.target_percent) is not int or not 1 <= self.target_percent <= 100:
10 raise ValueError("target_percent must be an integer from 1 to 100")
11 if type(self.saw_untrusted_instruction) is not bool:
12 raise ValueError("instruction flag must be Boolean")
13
14def required_reviews(proposal: PromotionProposal) -> set[str]:
15 reviews = {"RELEASE_APPROVAL"}
16 if proposal.saw_untrusted_instruction:
17 reviews.add("SECURITY_REVIEW")
18 if proposal.target_percent > 10:
19 reviews.add("EXCEPTION_REVIEW")
20 return reviews
21
22def review_state(proposal, decisions):
23 # Decisions come from authenticated reviewers in the real service.
24 needed = required_reviews(proposal)
25 if any(value not in {"APPROVED", "DENIED"} for value in decisions.values()):
26 raise ValueError("unknown review decision")
27 if any(decisions.get(name) == "DENIED" for name in needed):
28 return "DENIED_APPEAL_AVAILABLE"
29 if not needed <= decisions.keys():
30 return "HELD_MISSING_REVIEW"
31 return "APPROVED_FOR_EXECUTOR"
32
33proposal = PromotionProposal(target_percent=10, saw_untrusted_instruction=True)
34print("required reviews:", sorted(required_reviews(proposal)))
35print("before review:", review_state(proposal, {}))
36print("after denial:", review_state(proposal, {"SECURITY_REVIEW": "DENIED"}))
37print("after approval:", review_state(proposal, {
38 "RELEASE_APPROVAL": "APPROVED", "SECURITY_REVIEW": "APPROVED",
39}))1required reviews: ['RELEASE_APPROVAL', 'SECURITY_REVIEW']
2before review: HELD_MISSING_REVIEW
3after denial: DENIED_APPEAL_AVAILABLE
4after approval: APPROVED_FOR_EXECUTORAPPROVED_FOR_EXECUTOR isn't an executed effect. A real executor must bind approvals to the active operation and exact candidate, environment, traffic target, policy, and evidence versions, verify reviewer authority and freshness, and recheck before committing. The saw_untrusted_instruction flag is a fixture, not a reliable detector: server-side authorization remains necessary even when no attack is recognized. For a legally high-risk workflow, map oversight requirements to the applicable duties and roles.[8]
When is "human oversight" meaningful?
Answer
When a named reviewer can see the relevant evidence before a consequential effect, has authority to stop or change it, records a reason, and an affected person has an escalation path.
Gate releases on an evidence package
All these artifacts matter only when the deployment pipeline reads them. A release candidate can carry a folder full of documents and still ship an unsafe effect if no executable check consumes them. Make the candidate fail when required evidence is missing or a red-team finding remains open.
For the model-promotion assistant, the minimal package answers eight review questions:
| Evidence | Why it exists |
|---|---|
| Workflow memo | Records intended use, people affected, and dated review route |
| Risk row | Connects harm to control, owner, residual risk, and next review |
| System card | Identifies shipped policy, model, tools, limitations, and tests |
| Dataset record | Makes evaluation cases and labels reproducible |
| Audit replay | Shows a proposed promotion was gated and preserved correctly |
| Red-team report | Blocks release when an unsafe effect remains |
| Accessibility and slice report | Finds valid requests blocked on supported paths |
| Oversight path | Shows approval and appeal behavior exists |
The final check takes reviewed artifact references, not a list of filenames. Each result must belong to this candidate and the exact deployment configuration. A version label can stay unchanged while someone adds a tool or replaces the policy.
Here the host hashes a small synthetic manifest and compares that digest with each review. In production the manifest must bind all relevant immutable artifacts, including model, prompts, policy, tools, evaluation inputs, and environment. Reviews and finding counts must come from authenticated validators and reviewers bound to that manifest. This example assumes that provenance has already been verified; hashing self-reported claims wouldn't verify them. Passing makes the candidate eligible for release approval, without deploying it or establishing legal compliance.
1from dataclasses import dataclass
2import hashlib
3import json
4
5REQUIRED_EVIDENCE = {
6 "workflow_memo",
7 "risk_register_row",
8 "system_card",
9 "dataset_record",
10 "audit_replay",
11 "red_team_report",
12 "accessibility_and_slice_report",
13 "oversight_runbook",
14}
15
16@dataclass(frozen=True)
17class EvidenceReview:
18 candidate: str
19 deployment_digest: str
20 artifact_ref: str
21 accepted: bool
22
23def manifest_digest(manifest: dict[str, object]) -> str:
24 # One local Python encoding. Production manifests need a specified format.
25 payload = json.dumps(manifest, sort_keys=True).encode("utf-8")
26 return hashlib.sha256(payload).hexdigest()
27
28def release_decision(
29 candidate: str,
30 deployment_digest: str,
31 evidence: dict[str, EvidenceReview],
32 unsafe_effects: int,
33 open_accessibility_findings: int,
34 legal_signoff: bool,
35) -> tuple[bool, list[str]]:
36 if not isinstance(candidate, str) or not candidate.strip():
37 raise ValueError("candidate identity required")
38 if not isinstance(deployment_digest, str) or len(deployment_digest) != 64 or any(
39 char not in "0123456789abcdef" for char in deployment_digest
40 ):
41 raise ValueError("host deployment digest must be 64 lowercase hexadecimal characters")
42 if any(type(n) is not int or n < 0 for n in (unsafe_effects, open_accessibility_findings)):
43 raise ValueError("finding counts must be nonnegative integers")
44 if type(legal_signoff) is not bool:
45 raise ValueError("legal signoff must be Boolean")
46 reasons: list[str] = []
47 missing = sorted(REQUIRED_EVIDENCE - evidence.keys())
48 if missing:
49 reasons.append(f"missing evidence: {', '.join(missing)}")
50 for name in sorted(REQUIRED_EVIDENCE & evidence.keys()):
51 item = evidence[name]
52 if (
53 item.candidate != candidate or item.deployment_digest != deployment_digest
54 or not isinstance(item.artifact_ref, str) or not item.artifact_ref.strip()
55 or item.accepted is not True
56 ):
57 reasons.append(f"unaccepted or mismatched evidence: {name}")
58 if not legal_signoff:
59 reasons.append("required legal signoff missing")
60 if unsafe_effects:
61 reasons.append(f"unsafe effects remain: {unsafe_effects}")
62 if open_accessibility_findings:
63 reasons.append(f"accessibility findings remain: {open_accessibility_findings}")
64 return not reasons, reasons
65
66candidate = "promotion-assistant-7.2"
67deployment = {"model": "mock-model-v1", "prompt": "release-prompt-v3",
68 "policy": "promotion-gate-v7", "tools": ["propose_promotion"],
69 "eval": "frozen-promotion-eval-v2", "environment": "mock-prod"}
70expected_digest = manifest_digest(deployment)
71reviewed = {name: EvidenceReview(candidate, expected_digest, f"reviewed-artifacts/{name}/v7", True)
72 for name in REQUIRED_EVIDENCE}
73draft_evidence = {name: item for name, item in reviewed.items() if name != "dataset_record"}
74draft_ready, draft_reasons = release_decision(
75 candidate,
76 expected_digest,
77 draft_evidence,
78 unsafe_effects=1,
79 open_accessibility_findings=1,
80 legal_signoff=True,
81)
82print("draft_ready:", draft_ready)
83print("draft_reasons:", draft_reasons)
84
85reviewed_ready, reviewed_reasons = release_decision(
86 candidate,
87 expected_digest,
88 reviewed,
89 unsafe_effects=0,
90 open_accessibility_findings=0,
91 legal_signoff=True,
92)
93print("reviewed_ready:", reviewed_ready)
94print("reviewed_reasons:", reviewed_reasons)
95
96changed = {**deployment, "policy": "promotion-gate-v8"}
97changed_ready, _ = release_decision(
98 candidate, manifest_digest(changed), reviewed,
99 unsafe_effects=0, open_accessibility_findings=0, legal_signoff=True,
100)
101print("changed_configuration_ready:", changed_ready)1draft_ready: False
2draft_reasons: ['missing evidence: dataset_record', 'unsafe effects remain: 1', 'accessibility findings remain: 1']
3reviewed_ready: True
4reviewed_reasons: []
5changed_configuration_ready: Falsedataset_record is release evidence, not administrative decoration. Without its provenance, labels, split, and retention rule, a reviewer can't tell where the score came from or reproduce it safely.
The last line keeps the candidate name but changes the policy. Its new manifest digest rejects the old reviews. Try adding a deployment tool: which evidence must be rerun, and which owner decides? Invalidate affected reviews and preserve the previous decision record. The artifact list and thresholds belong to this workflow; other systems need a different package based on their harms, users, and duties.