Teams using GitHub Copilot reported a 55% increase in code merged per developer in Microsoft's 2023 study, and incident rates at those same organizations went up. That gap between output velocity and actual quality is the measurement problem every staff engineer needs to solve before their org mistakes throughput for health.
The metric explosion following AI adoption has been predictable and mostly useless. Dashboards fill up with lines of code accepted, Copilot suggestion acceptance rates, and PR velocity. None of those numbers have a reliable causal path to fewer production incidents, faster code reviews, or reduced technical debt. Measuring them is worse than measuring nothing, because they create a false floor of confidence that delays the harder conversation.
Why Acceptance Rate Is a Trap
Suggestion acceptance rate measures how often a developer presses Tab. It does not measure whether pressing Tab was correct. A developer who accepts 90% of Copilot suggestions on a greenfield CRUD service is making a different decision than one accepting 90% of suggestions in a Kafka consumer that handles exactly-once delivery guarantees. Aggregating those into one number and calling it "AI adoption" is the kind of metric that sounds rigorous in a board slide and falls apart the first time an SRE opens a postmortem.
The deeper problem is selection bias in the tooling itself. Copilot surfaces confident-looking suggestions regardless of whether the surrounding context supports them. Developers who are less experienced in a domain accept more suggestions, not because the suggestions are better, but because they lack the pattern recognition to reject them. High acceptance rate in a domain where your team is thin on expertise is a warning sign, not a success metric.
Acceptance rate also collapses across suggestion types. Completing a for loop variable name is categorically different from autocompleting a SQL query that touches a billing table. Any aggregated acceptance metric that does not separate these is not measuring AI code quality. It is measuring how often a developer's hands stayed still for 200 milliseconds.
The Metrics That Actually Correlate
Post-merge review churn is the first metric worth tracking at the org level. Measure the number of comments added to a PR after it is already merged, surfaced through your code review tooling, and broken out by whether the PR was AI-assisted. Specifically, you want review_comment_velocity_postmerge normalized per thousand lines. If AI-assisted PRs are generating 40% more post-merge comments than human-only PRs, something is failing in the review step, not the writing step.
Here is a rough Datadog custom metric query structure to get you started:
-- Normalized post-merge churn by PR type
SELECT
pr_type, -- 'ai_assisted' | 'human'
COUNT(review_comments.id)
/ NULLIF(SUM(pr_diff_lines), 0)
* 1000 AS comments_per_kloc,
AVG(EXTRACT(EPOCH FROM
(comment_created_at - pr_merged_at))
/ 3600) AS avg_hours_after_merge
FROM pull_requests
JOIN review_comments USING (pr_id)
WHERE pr_merged_at >= NOW() - INTERVAL '30 days'
AND comment_created_at > pr_merged_at
GROUP BY pr_type;
You need this from your version control platform's API, not from Copilot's dashboard. GitHub, GitLab, and Bitbucket all expose enough webhook data to build this. If post-merge churn is trending up alongside AI adoption, that is your signal to look at review process design, not at the AI tooling itself.
Incident attribution by change origin is the second metric, and it is harder to implement but more honest. Every post-incident review should tag the offending change with whether it was AI-assisted, human, or mixed. Over a 90-day window you want the ratio of incidents-per-change-origin against the volume of changes from each origin. This is not about blaming AI. It is about understanding whether your review gates are calibrated correctly for the new input source.
Most orgs skip this because their incident tooling (PagerDuty, Opsgenie, Jira) is not connected to their deployment pipeline metadata. That gap is solvable with a build tag and about two days of integration work. If you are not doing it, you are flying blind on the most important quality signal you have.
Semantic drift rate is the third, and the most underused. Run a static analysis pass on AI-assisted PRs against your internal style guide and architecture decision records. What percentage of AI-assisted changes introduce patterns that are locally coherent but globally inconsistent with your codebase? Things like using a new HTTP client library when the org standard is already pinned, or introducing a new error-handling pattern in a service that has an established one.
This one requires investment. You need to encode your architectural decisions as machine-readable rules, which is work most teams have not done. It is also the metric that most directly measures whether AI is accelerating consistency or eroding it. OpenThunder is built around exactly this kind of architecture-aware verification, running checks against your actual system constraints before changes ship rather than after.
How to Implement This Without Burning Political Capital
The organizational challenge with these metrics is not technical. It is that they produce uncomfortable findings. Post-merge churn data will show that some teams are rubber-stamping AI output. Incident attribution will show that certain service areas have poor review hygiene regardless of origin. Semantic drift analysis will show that your architecture decisions are not actually enforced anywhere.
Roll these out in phases. Start with post-merge churn because it is the easiest to explain, the least politically charged, and the most immediately actionable. A team lead can look at that number and immediately understand what lever to pull: slow down reviews, add a checklist, require a second reviewer on AI-assisted PRs above a certain diff size. You do not need to redesign the org to respond to it.
Incident attribution is the second phase, and it requires buy-in from your SRE or platform team. Frame it as improving your postmortem fidelity, not as auditing AI. Most SREs are already frustrated by the loss of signal in postmortems as change velocity increases. Giving them a new attribution dimension is usually welcome.
Semantic drift analysis is phase three and requires staff or principal-level sponsorship. You are asking engineering to codify its own architectural decisions, which means you are also forcing a conversation about which decisions are actually decisions and which ones are vibes that have never been written down. That is valuable work. It is also slow and occasionally contentious.
These metrics only work if you measure them before AI adoption fully saturates your team. If you wait until 80% of code is AI-assisted, you lose your baseline. Establish the measurement infrastructure now, even if the numbers look fine today.
For teams that want to move faster on the semantic drift problem specifically, the approach OpenThunder takes is worth understanding: rather than post-hoc analysis, it runs verification inline as part of the merge pipeline, which shifts the enforcement point left and removes the need for a separate audit loop.
The Org-Level Framing That Changes the Conversation
The reason most AI quality metrics fail is that they are framed as developer productivity metrics. They measure individuals and aggregate up. The org-level metrics described here are different: they measure the system's response to AI-generated input, not the input itself.
This framing matters because it correctly locates accountability. If post-merge churn is high, the review process is the intervention point. If incident attribution shows AI-assisted changes failing at higher rates in one service area, the review culture in that area is the intervention point. The AI tooling is the constant. Your processes are the variable.
Senior engineers often get asked to evaluate whether the org should adopt AI coding tools, and they answer the wrong question. The question is not whether to adopt. The question is whether your quality infrastructure can absorb the change-volume increase without proportional degradation. Most orgs find out the answer is no, but only after the incidents.
Measure the system. Tune the system. The AI will keep generating code either way.
See what OpenThunder verifies
OpenThunder independently verifies AI-assisted changes against your architecture, security, and intent before they ship, giving your review process the signal it needs at the speed AI adoption demands. Try it here.
The metric you choose to optimize is the org you end up with: pick throughput, and you will optimize throughput straight into the next outage.