Operating model and RACI for governance
Trust Architecture Playbook: Governance pillar
Central vs. delegated responsibilities
Centralized control and delegated execution are both required. Over-centralization creates bottlenecks and encourages shadow PKI. Over-delegation without guardrails creates issuance sprawl and unaccountable trust. The operating model should centralize policy, profiles, CA sources, and exception governance while delegating execution to teams that own the platforms and services - within strict boundaries.
Operating area | Central PKI / Security responsibilities | Delegated team responsibilities |
|---|---|---|
Policy and trust architecture | Define trust domains, PKI policy, CA source registry, profile standards, and exception criteria. | Use approved trust domains and profiles; request changes through governance. |
Profiles and templates | Approve profile catalog, constraints, business-unit mapping, CA source pairing, and retirement. | Request profiles with documented use cases; operate within approved constraints. |
Identity and access | Define roles, service-user standards, delegation boundaries, and access review cadence. | Protect credentials, request scoped access, rotate secrets, report ownership changes. |
CA sources and connectors | Approve CA sources, connector patterns, root/intermediate registration, and decommission plans. | Maintain platform prerequisites, sensor/connector reachability, and operational health. |
Lifecycle operations | Define guardrails, evidence requirements, break-glass, revocation policy, and audit reporting. | Execute renewals, deployments, validation, incident response, and owner updates. |
Audit and metrics | Define metrics, evidence packs, reporting cadence, and audit support. | Remediate findings, close exceptions, and provide platform/service context. |
Role-specific accountability in plain language
Wichtig
Accountability rule
Central teams own the guardrails. Delegated teams own execution within those guardrails. Service owners own business risk. GRC verifies evidence. Executives approve residual risk when the standard control model cannot be met.
The RACI defines formal accountability. The following narrative is the simpler operating contract most teams need day to day.
Role | What they own | What they provide |
|---|---|---|
Central PKI / Security | PKI policy, trust-domain registry, CA source registry, profile catalog, service-user rules, access model, exception criteria, evidence standards, and crypto-agility roadmap. | Approved guardrails, profile decisions, access reviews, CA-source approvals, exception decisions, and executive-ready risk reporting. |
Platform teams | Connectors, sensors, agents, deployment patterns, platform credentials, automation identities, validation paths, reload behavior, and operational runbooks for the platforms they operate. | Deterministic deployment, monitoring, failure remediation, credential rotation, platform health, and evidence of successful deployment or recovery. |
Service owners | The business service, namespace entitlement, production change impact, criticality tier, certificate ownership, approval of material changes, and acceptance of service risk. | Accurate owner metadata, service context, escalation path, approval for production use, validation participation, and remediation prioritization. |
GRC / Audit | Control mapping, evidence expectations, audit sampling, reporting requirements, and independent review of whether the operating model is producing defensible evidence. | Audit criteria, evidence-pack sampling, findings, remediation tracking, and assurance reporting. |
Executives / Risk owners | Risk appetite, funding priority, residual risk acceptance, escalated exception review when board-level risk acceptance exceeds standard tolerance, and escalation when remediation conflicts with business priorities. | Risk decisions, resourcing, deadline enforcement, and acceptance or rejection of deviations that exceed operational tolerance. |
Service owners
Service owners are accountable for the business and operational context of the certificates protecting their services.
Service owners do not need to become PKI engineers, but they do need to own the business and operational context of the certificates that protect their services. The service owner is the person or team that can say whether a certificate is expected, where it is used, what happens if it fails, and whether a requested deviation is worth the risk.
Expectation | What the service owner does | Evidence or signal |
|---|---|---|
Own the service context | Confirm service name, business owner, technical owner, criticality, exposure, environment, and support path. | Inventory owner fields, mandatory metadata, and escalation record. |
Authorize namespace use | Approve the domain, hostname, SAN, URI SAN, service account, or workload identity used by the service. | Domain/namespace registry entry and request approval. |
Accept lifecycle impact | Understand renewal, rekey, revocation, retirement, and automation changes that can affect availability. | Lifecycle approval, change record, and validation evidence. |
Support remediation | Resolve missing ownership, metadata gaps, manual deployment blockers, shared certificate risk, or exception retirement. | Remediation ticket, closure evidence, and dashboard trend. |
Participate in exceptions | Own the business risk when a standard control cannot be met and commit to a retirement plan. | Exception register entry with risk owner, technical owner, expiry, and compensating controls. |
Before approving a certificate request, lifecycle change, or exception, a service owner should be able to answer these questions without relying on the PKI team to infer business context.
Question | Why it matters |
|---|---|
What service or product does this certificate protect? | Connects the certificate to business impact, incident routing, ownership, and funding priority. |
Are the requested DNS names, URI SANs, service accounts, or workload identities legitimately part of this service? | Prevents namespace drift and unauthorized identity binding. |
What happens if this certificate expires, is revoked, or is replaced incorrectly? | Sets the criticality tier, approval path, monitoring, and rollback expectation. |
Is this a routine lifecycle action or a material change? | Distinguishes normal renewal from SAN expansion, CA migration, profile change, algorithm change, or exception-driven issuance. |
What evidence will prove the service is presenting the intended certificate after the change? | Closes the loop between issuance, deployment, validation, and audit evidence. |
Platform owners
Platform owners are accountable for the certificate execution surface their platforms operate.
Platform owners are accountable for the certificate execution surface: the connectors, agents, sensors, ingress controllers, load balancers, vaults, service meshes, cluster issuers, reload behavior, validation hooks, and operational runbooks that make governed lifecycle work reliable.
Expectation | What the platform owner does | Evidence or signal |
|---|---|---|
Register the platform boundary | Document platform, environment, business unit, namespace scope, connector pattern, and supported certificate types. | Platform registry, CA source registry linkage, and approved profiles. |
Maintain integration health | Operate connectors, sensors, agents, cert-manager, vault integration, cloud certificate services, or load balancer integrations. | Connector health, scan freshness, deployment success, and failure trends. |
Protect automation identities | Use scoped service users, ACME accounts, connector credentials, and secret storage with rotation and emergency disable paths. | Service-user register, access review, credential rotation record. |
Publish lifecycle runbooks | Define install, bind, reload, validate, rollback, revocation, and break-glass steps for each platform pattern. | Runbook, tabletop result, and validation evidence. |
Close platform gaps | Escalate repeated exceptions as pattern gaps, not one-off service problems. | Pattern backlog, exception clustering report, and remediation plan. |
Before a platform becomes a governed lifecycle path, the platform owner should be able to prove that the execution surface is deterministic, monitored, and recoverable.
Question | Why it matters |
|---|---|
Which connector, agent, sensor, ACME client, issuer, or API pattern owns this platform? | Prevents unmanaged issuance paths and makes support ownership explicit. |
Which service users, credentials, or workload identities can request or deploy certificates? | Defines delegated issuance authority and the scope of credential rotation and disablement. |
How does the platform install, bind, reload, and validate certificates? | Separates issuance success from verified deployment. |
What is the rollback, reissue, revocation, or break-glass path for high-criticality services? | Makes recovery an engineered path rather than an incident-time improvisation. |
Which repeated exceptions are really platform pattern gaps? | Turns clusters of one-off exceptions into backlog items that improve the standard model. |
Responsible, Accountable, Consulted, Informed (RACI)
Activity | PKI governance board | Central PKI / Security | Platform / Infra | App / Service owner | Compliance / GRC |
|---|---|---|---|---|---|
Approve PKI policy and trust domains | A | R | C | C | C |
Approve CA source | A | R | C | I | C |
Approve profile standard | A | R | C | C | C |
Create / maintain profile | C | R/A1 | C | C | I |
Approve delegated RA scope | A | R/A1 | C | C | C |
Operate connectors / sensors / agents | I | C | R/A | C | I |
Request certificate | I | C | C | R/A | I |
Maintain certificate metadata and owner mapping | I | C | R | R/A | I |
Execute renewal/deployment validation | I | C | R/A | R | I |
Approve Tier 0/1 exception | A | R | C | C | C |
Audit evidence and reporting | I | C | I | I | R/A |
Break-glass invocation | I | R/A | R | C | I |
Break-glass post-review | A | R | C | C | C |
| |||||
Governance cadence
Cadence | Governance activity | Outputs |
|---|---|---|
Weekly | Triage unowned certificates, CT anomalies, expiring high-criticality certificates, and urgent exceptions. | Remediation assignments, incident escalations, owner updates. |
Monthly | Review high-risk profiles, Tier 0/1 lifecycle failures, break-glass events, and open exceptions. | Corrective actions, profile changes, exception decisions. |
Quarterly | Access review, CA source operational review, profile catalog review, delegated RA review, metrics review. | Access removals, profile retirements, risk dashboard, governance decisions. |
Semiannual | Crypto-agility exercise, mass-rotation tabletop, CA distrust simulation, evidence pack sampling. | Exercise results, readiness score, remediation backlog. |
Annual | PKI policy review, trust-domain registry review, CP/CPS alignment, audit support package, CA source formal re-approval. | Approved policy updates, risk acceptance refresh, audit evidence set, renewed CA source approvals. |