Architecture governance fails at the document layer roughly 90% of the time. The gap is not that engineers write bad ADRs. It is that no runtime mechanism connects the document to the code that ships.
Why Architecture Decision Records Stop Working After Team Six
Architecture Decision Records were designed to capture the why behind a structural choice, not to enforce it. That distinction matters enormously and most teams miss it entirely.
When a team is five or six engineers, ADRs work fine as a social contract. Everyone has read the same documents, everyone was in the same Slack thread, and tribal memory fills the gaps. Scale to thirty engineers and two ADRs are added per month while four engineers who have never read the backlog join every sprint. The ADR lives in Confluence or a /docs folder. Nobody links to it from the PR template. Nobody checks whether the PR violates it. The document is not governance; it is archaeology.
The specific failure mode is drift by omission. A new service gets wired to a legacy synchronous HTTP endpoint instead of the approved Kafka topic because the engineer building it did not know the decision existed. The violation is not malicious. It is invisible. By the time someone notices, three downstream services have taken the same shortcut, and the pattern is de facto policy.
A decision record has no enforcement surface. It cannot reject a pull request. It cannot fail a build. It cannot alert on a deployment.
The Three Layers Where Policy Actually Lives
Effective architecture governance operates at exactly three layers. Most organizations only cover one or two.
The first layer is static analysis at commit time. Tools like ArchUnit for Java, Dependency Cruiser for Node.js, or custom Semgrep rules for cross-language enforcement run in CI and fail the build when a module boundary is crossed illegally. A service in the payments domain importing directly from inventory internals is a build error, not a code review comment. This is the highest-leverage layer because it is immediate, cheap, and machine-enforced.
The second layer is infrastructure policy as code. This is where Open Policy Agent (OPA) and tools like Checkov or Terraform Sentinel live. If your architecture decision says all external-facing services must route through the API gateway, that invariant lives as a Rego policy that runs against every Terraform plan before terraform apply is allowed to proceed. The gateway bypass cannot be deployed by accident; it requires a deliberate policy override with a documented exception.
The third layer is runtime observability with policy coupling. This is the most underused layer. Service mesh configurations in Istio or Linkerd can enforce that service A is not allowed to call service C directly. Datadog or Honeycomb trace maps can detect emerging call patterns that violate the intended topology and alert before they calcify. This layer catches the violations static analysis cannot see: runtime wiring via environment variables, dynamic service discovery, or hand-rolled HTTP clients that do not import anything policy-checked.
Most teams have partial layer one coverage and almost no layer two or three. That leaves the biggest blast radius categories completely unguarded.
Policy as Code: What to Actually Write
The phrase "policy as code" gets thrown around loosely. Here is what a useful policy looks like versus a useless one.
A useless policy checks formatting or naming conventions that a linter already covers. If your OPA policy is asserting that Terraform resource names use underscores instead of hyphens, you have not added governance; you have added friction with a fancier tool.
A useful policy encodes a structural decision that has real security, reliability, or cost consequences if violated. Examples worth writing:
- Every S3 bucket that stores PII must have a specific tag that ties it to a data classification record. The policy fails any
aws_s3_bucketresource missing that tag. - No Lambda function in the
paymentsnamespace may have a VPC configuration that routes to a subnet outside the PCI-scoped CIDR range. - No service in the
reportingdomain may declare a synchronous dependency on a service in theorder-processingdomain. Async event subscriptions are permitted; direct HTTP calls are not.
The pattern in every useful policy is the same: a structural invariant derived from a real architecture decision, expressed as a machine-checkable predicate. The ADR still exists and explains why the invariant is what it is. The policy is what actually enforces it.
Writing these requires a conversation between the staff or principal engineer who owns the decision and whoever owns the CI/CD platform. That conversation almost never happens organically. It needs to be a formal step in the ADR process: every ADR with enforcement consequences gets a linked policy PR. No policy PR, no merged ADR.
The Organizational Work Static Analysis Cannot Do
Policy as code handles the mechanical enforcement layer. Architecture governance also has a human-coordination layer that tools cannot replace.
The failure mode here is what I call exception accumulation. A team ships a policy violation under a documented exception. The exception is reasonable: there is a deadline, the approved pattern is not ready, the tech debt is acknowledged. Six months later there are forty exceptions and the policy is effectively dead. Nobody revoked it, but it no longer means anything because the exception rate is 80%.
The fix is a structured exception review cadence. Every active policy exception gets a quarterly review in the staff engineer or architecture review forum. Three questions: Is the underlying blocker that justified the exception still real? Has the exception spawned copies elsewhere? What is the migration path to compliance? Exceptions that survive two quarterly reviews without a concrete migration plan get escalated to a decision: either extend the policy to formally permit the pattern, or commit engineering time to retire the exception.
This sounds bureaucratic. It is, deliberately. The bureaucracy is the feature. Architecture drift happens when exceptions are cheap and invisible. Making exceptions slightly expensive and highly visible is how you keep the policy meaningful.
There is a separate staffing consideration. Policy-as-code infrastructure needs an owner. At a fifty-engineer company, this is probably a staff engineer with a twenty-percent allocation. At a three-hundred-engineer company, this is a platform team function. If nobody owns the policy infrastructure, policies go stale as the codebase evolves, and a stale policy that blocks legitimate work is worse than no policy at all because it trains engineers to route around the system.
Where AI-Assisted Development Changes the Calculus
The arrival of AI coding assistants has not changed the architecture governance problem structurally. It has amplified it by roughly an order of magnitude in velocity terms.
A senior engineer working manually might introduce one architecture violation per week through oversight. A team using Copilot or Claude to generate scaffolding can introduce ten, because the model has no knowledge of your specific structural decisions. It knows common patterns. It does not know that your team decided synchronous cross-domain calls are prohibited, or that the legacy auth service is off-limits to new integrations.
This is where the enforcement layers become critical rather than merely useful. Static analysis and OPA policies do not care whether the code was written by a human or generated by a model. The PR fails either way. That consistency is the whole point.
OpenThunder was built specifically for this gap: verifying that AI-assisted changes conform to the architecture, security, and intent constraints your team has actually defined, before those changes reach the review queue. The model generating the code does not know your architectural invariants. OpenThunder does.
The broader principle is that your governance infrastructure needs to be calibrated for the actual throughput of code changes in your system, not the throughput you had two years ago. If AI assistance has tripled the number of PRs your team opens, your policy enforcement surface needs to scale with that. An ADR that two people read per month is not scaling with anything.
Putting It Together: A Governance Stack That Actually Holds
A governance architecture that holds under team growth and AI-assisted development looks like this:
The ADR is the source of truth for the decision. It captures context, alternatives considered, and consequences. It links to the policies that enforce it. Without that link, the ADR is documentation; with the link, it is the head of an enforcement chain.
Static analysis tools run in CI at the module and dependency level. ArchUnit, Dependency Cruiser, or Semgrep custom rules are the right instruments depending on language stack. These fail builds, not code reviews.
OPA or Sentinel policies enforce infrastructure invariants. Every architecture decision that has an infrastructure expression lives as a machine-checkable policy in the IaC pipeline. No exceptions without a ticket, no ticket without a quarterly review.
Runtime observability closes the gap. Service mesh policies and distributed trace anomaly detection catch the violations that happen outside the static analysis surface. Datadog's service map or Honeycomb's trace explorer are your audit trail.
Exception management is a first-class process. Every exception is logged, linked to an owner, and reviewed quarterly. High exception rates trigger policy review, not policy abandonment.
This stack is not free to build. The upfront cost is real, probably four to eight weeks of platform-focused engineering to wire up the first complete pass. The ongoing cost is maintenance, roughly proportional to the rate at which your architecture evolves. The alternative cost, the one you pay in production incidents, security findings, and architectural spaghetti, compounds silently and is far larger by the time it surfaces.
OpenThunder fits into this stack at the AI-velocity layer, catching what static analysis misses when the code generation rate outpaces human review capacity. No single tool covers the full surface. Governance at scale is a stack, not a solution.
If your architecture policy only exists in a document, you do not have a policy. You have a suggestion.
See what OpenThunder verifies
OpenThunder independently verifies AI-assisted changes against your architecture, security, and intent before they ship. If you are building out a governance stack and want a layer that understands your specific structural constraints rather than just generic best practices, try it here.
The only architecture governance that works is the kind the codebase cannot ignore.