Part 4 of 4: Rethinking the quality operating model for AI-native engineering.
Imagine this scenario:
Two teams ship the same product on the same day. Team A ran a suite of tests via AI agents, got green across the board, and shipped. Team B ran the same suite of tests via AI agents, got the same green results, but can also tell you what was tested, how it was tested, what the agents decided to focus on and why, what risks were assessed and accepted, and who made those calls. Both teams shipped but only one team can really stand behind what they’ve delivered.
The distinction I’m making is whether the team has earned confidence in what the testing tells them, or assumed confidence that hasn’t really been tested, and that confidence shouldn’t be thought of as simply whether testing happened. I think that distinction (earned vs assumed) is the most useful frame for thinking about governance in an AI-native quality system.
Governance, in the context of everything I’ve covered in this article series, is the set of practices that turn assumed confidence into earned confidence, is what makes speed sustainable rather than reckless, and is the layer that makes the operating model from article 3 something you can actually defend, not just describe.
And this isn’t a set of approval gates or a compliance exercise I’m talking about. Those framings are what give governance its reputation for slowing things down. Designed properly, governance is what actually enables you to move fast with confidence, rather than move fast and hope for the best.
Why the old governance model doesn’t hold up
Without AI-assisted testing, governance is largely a “human discipline” question. What testing was completed, and by whom? Did someone review the results and use the information to make an informed decision? The evidence trail is whatever someone produced and recorded. The failure modes are related to people being inconsistent, there being coverage gaps, and a ceiling on speed that overbloated processes and approval work imposes. It’s imperfect, but you can audit it – the people are visible in the process.
With AI-assisted testing, the governance surface changes shape in ways that most teams haven’t fully thought of yet. The AI is doing more of the work, generating and executing tests, making coverage decisions, interpreting behaviours, identifying risks, discovering defects and logging them within the existing tools already in play. That’s more work that’s happening faster. But it means more decisions being made that nobody explicitly authorised with a timestamped audit trail with their name against it when doing so. It also means there’s more coverage that might exist on paper but not in our described practices, and more findings that got acted on or dismissed without a clear human decision point attached to them.
The specific risk that emerges is different in kind from human testing failures. It’s coverage you think exists but doesn’t, or findings you trusted but shouldn’t have, or context the agent was missing that nobody noticed because the agent was confident anyway… These aren’t the same failure modes as a tester who ran out of time, so they need different responses in terms of governance.
And there’s one more dimension I want to name here too: the more AI is involved on either side of the testing equation, whether AI is part of the product being built or its AI tools being used within the testing process, the more these governance questions compound. The principles in this article apply whether your product has an AI in it or not, but the stakes get higher as AI involvement increases on either side, and the “assumed confidence” trap gets harder to spot.
The four governance concerns
I find it useful to think about governance sitting across four interconnected concerns: evaluation, evidence, confidence, and risk. Miss any one of them and the confidence the system produces is hugely affected.
Evaluation
How do you know the system is testing the right things, in the right ways, at the right depth?
This term “evaluation” has been used a lot recently, especially in the AI context, and with prominent CEOs of huge AI tech companies talking about the need for slowing down to be able to do more “evaluation”. And also in the quality communities too, with more people talking about the need to build AI Evaluation Frameworks… The word has been used for decades though! Fun fact: my first role in tech was as a “Graduate Evaluation Engineer”. The role? To “evaluate” products. So for me, the term has somewhat had a level of interchangeability with the word “testing” (but as we know, even the word “testing” has had its challenges regarding misunderstandings).
To me, evaluation is the ongoing question of whether the quality system’s judgement is actually good enough to trust. With human testing, evaluation is partly implicit, experienced testers develop a feel for whether their coverage is meaningful, whether they’re finding things that matter. With AI-assisted testing, evaluation needs to be more deliberate, because the system will confidently generate tests and report coverage whether or not that coverage is meaningful.
The governance practice here is building in regular moments to evaluate the evaluator. Reviewing what the agents chose to test and why. Comparing agent coverage against human risk judgement. Asking whether the things the system is focused on are actually the things the team is most worried about. And adjusting context and constraints when the answer is no.
This is an ongoing discipline, and it probably belongs in the feedback loop (Layer 5 of the operating model from article 3 in this series) rather than as a separate governance process bolted on the side. It really shouldn’t be thought of as a one-time setup activity.
Evidence
What does the system produce that you can actually point to?
Evidence is the record of what happened, what was tested, what was discovered, what was decided. Without AI-assisted testing, evidence is typically a mix of test reports, defect logs, and whatever notes a tester managed to write up, which is often less than anyone would like. With AI-assisted testing, the potential for richer, more structured evidence is genuinely significant, but only if the governance model is designed to capture and use it.
The governance question is whether that output is in a form that’s useful to someone who wasn’t in the room when decisions were made. It’s not whether the system produces output at all. It needs to be useful to the developer or tester who ran the test, but also to an engineering manager or team lead reviewing a production incident, a product leader making a risk call, or an engineering leader being asked to account for a release decision three months later.
Confidence
How does confidence in the system build over time, and how do you know when it’s genuine rather than assumed?
Confidence is the output that governance is ultimately trying to produce. Like the difference between a team that ships because the tests went green, and a team that ships because they understand what the tests covered, what they didn’t, what risks were accepted, and why.
Earned confidence also means confidence under pressure. The kind that holds up when something goes wrong in production and someone asks how it got through. Building it requires the trust-building practices from article 2, visible evidence (the second concern above), and the feedback loops from Layer 5 of the operating model. It compounds over time, but only if the other three concerns are in place.
Risk
Where is it most likely to go wrong, and what are you doing about it?
Risk governance in an AI-native quality system is less about eliminating risk and more about knowing where it sits and being honest about it. The old model surfaced risk through defect counts and test failures, lagging indicators that told you what went wrong after it went wrong. But like Quality Engineers have been trying to say for over a decade regarding risks that the AI-native model is surfacing now: risk is a leading signal – areas of the product with low coverage relative to change rate, agent findings that weren’t fully investigated, test debt accumulating in specific modules.
Governing risk means making those signals visible, making sure someone is reading and acting on them, and being honest about what the system doesn’t yet know. “We don’t have good coverage of this area and here’s what we’re doing about it” is a governance statement. “Everything’s green” when you haven’t checked whether “green” actually means anything is not.
What good governance actually looks like in practice
Rather than a framework, I want to give you four questions. These are the ones most teams haven’t explicitly answered, and they’re the ones that matter most when something goes wrong.
“What is the system authorised to do autonomously, and what needs a human decision?”
If this isn’t written down somewhere, it isn’t governed. The boundary between agentic and human-directed execution needs to be explicit, not implicit. Teams that haven’t had this conversation are making autonomy decisions by default, and defaults aren’t governance, they’re just whatever the tool does when nobody’s told it otherwise.
“What evidence does the system produce, in what form, who can access it, and for how long?”
The evidence needs to be in a form that’s useful to someone who wasn’t in the room when the decision was made. Not just a developer reviewing a test run, but a leader being asked to account for a release, or a team trying to understand a production incident. If the only person who can interpret the evidence is the person who generated it, the evidence isn’t doing its governance job.
“How does context get into the system, who owns it, and how does it stay current?”
Context quality is the new test plan (as I covered in Layer 1 of the operating model in article 3). Governance means someone owns it, someone updates it when the product changes or the risk landscape shifts, and there’s a clear process for ensuring the agents are testing against what actually matters right now, not what mattered three months ago.
“What happens when the system gets it wrong?”
Not if. When. Who is accountable? What’s the escalation path? How does a missed defect in a critical area get investigated, and what does the team learn from it that changes how the system is governed going forward? A quality system without a clear answer to this question is just running, and it shouldn’t be thought of as governed.
What leaders need to unlearn about governance
Three specific shifts, because the mental models most leaders carry about governance are almost perfectly calibrated to make it worse rather than better.
Governance is what enables speed, and it shouldn’t be thought of as what slows things down.
Teams that bolt governance on after the fact are the ones that slow down, because they’re constantly scrambling to produce evidence retrospectively, to reconstruct decisions that weren’t recorded, to explain coverage that was never properly evaluated. Teams that design governance into their system tend to move faster, because the evidence exists already, the confidence is earned as they go, and there’s no scrambling when someone asks a hard question.
Governance is a byproduct of good operating model design, and it shouldn’t be thought of as someone else’s job.
Legal can’t design your evidence trail for you. Compliance can’t define what the system is authorised to do. A QA manager somewhere can’t make the system’s coverage decisions legible. In an AI-native quality system, governance is a consequence of how the operating model is designed, and the operating model is the Engineering Leader’s responsibility. If you haven’t designed for evaluation, evidence, confidence, and risk, nobody else is going to do it for you.
You govern the system, and it shouldn’t be thought of as governing the tool.
You govern the system, and governing BearQ, or any agentic testing tool, shouldn’t be thought of as the same as governing the quality operating model it sits inside. The tool is one component. The governance model needs to cover what flows in, what the agents are authorised to do, what gets recorded, who reviews what, and what happens when something goes wrong. Scoping governance to the tool while leaving the system ungoverned is a bit like governing the car and ignoring the road.
BearQ as a governance-aware design
One of the things I noticed when using BearQ over a longer period is that several of its design choices aren’t just UX decisions or feature decisions. They’re governance decisions. And I can name them explicitly rather than leaving that implicit.
The most significant is the human-in-the-loop architecture, which gives teams three specific controls: focus areas, guardrails, and work reviews. Together, these three things give teams genuine control over where the dial sits between autonomous and directed operation, and crucially, that dial can move over time as trust and confidence build.
Focus areas determine where the agents direct their attention, which parts of the application, which behaviours, which risk areas. Setting focus areas is the mechanism for helping to guide the system towards what actually matters to the team rather than making arbitrary coverage decisions. It’s also a forcing function for a conversation most teams haven’t had explicitly: what do we actually consider highest risk right now? Having to answer that question in order to configure the system is itself a governance practice.
Guardrails define what the agents are and aren’t authorised to do. This is the explicit boundary between agentic and human-directed execution that the autonomy governance question demands. Teams that have set guardrails have answered “what is the system authorised to do autonomously?” Teams that haven’t are governing by default. Guardrails turn an implicit assumption into an explicit decision, which is exactly what governance is supposed to do.
Work reviews are the human checkpoint, the moment where a person evaluates what the agents proposed or produced before it’s acted on. This is where the evidence and confidence concerns connect to human judgement. A work review is a trust-building mechanism that lets confidence in the system grow incrementally, and it shouldn’t be thought of as just a quality gate. The team can see what the agent did, evaluate whether it made good decisions, and build a genuine picture of the system’s reliability over time rather than having to take it on faith.
The governance significance of all three together is that they make the autonomy level a deliberate, visible choice rather than a default. A team that starts with tight guardrails and frequent work reviews can extend autonomy gradually as confidence builds, and they can do it on their own terms rather than by accident. That’s governance built into the architecture, and it shouldn’t be thought of as just good product design.
Something I want to be clear about: BearQ gives you the raw material for governance. The governance model itself, the four questions, the four concerns, the decisions about what gets recorded and who reviews what, that’s still something that leaders have to design. The tool makes it possible but the operating model is actually what makes it real.
You can try BearQ now, on the 7-day free trial.
Closing: standing behind what you build
I want to close this series with a genuine sentiment, having spent the last few months using these tools, writing about them, and thinking about what this all means for the engineering leaders I know and work with.
The organisations that lead on quality over the next few years probably won’t be the ones that adopt AI testing tools the fastest. They’ll be the ones that build quality systems they can actually stand behind. Systems where confidence is earned rather than assumed, where evidence exists and is legible, where risk is visible and honestly accounted for, and where the humans in the loop are there because their judgement matters.
That’s what SmartBear calls “application integrity”, which I think is a useful way to name what I’ve been building towards across this whole series. It’s not a compliance checklist or an audit trail, but is the continuous, measurable integrity of the software, with the governance to operate at the pace that companies demand with AI-native engineering. That’s what “earned confidence” looks like at the system level in my opinion, and it’s what the four concerns in this article (evaluation, evidence, confidence and risk) are all trying to produce.
That’s what this series has been trying to map. Quality as a distributed system that needs to be designed deliberately. Trust as the human bottleneck that needs to be built intentionally. An operating model as the blueprint that connects the two. And governance as the layer that makes all of it something you can defend (not just describe).
The leaders who engage with these questions now, rather than waiting for a governance framework to be handed down from above, are the ones who will shape what good looks like. And the teams working for those leaders will be the ones who know the difference between shipping because the tests went green, and shipping because they actually understand what that means.
Those are very different things to be able to say, and the gap between them is ultimately what governance is for.
Thanks to SmartBear for the space to think through all of this properly, and for giving me early access to BearQ to test these ideas against my own web app.