From Theory to Practice: What an AI-Native Quality Operating Model Actually Looks Like

Part 3 of 4: Rethinking the quality operating model for AI-native engineering

If you’ve been following this series so far, you’ll know in the previous two articles, I’ve covered: why quality has outgrown the old operating model, and why trust is the real bottleneck when teams try to adopt AI testing tools. Being completely honest, both of those articles were more based on my diagnosis of the situations I see and hear. I tried to name the problem clearly, but I didn’t answer the question I suspect most people were asking and sitting with since reading them… “OK, Dan… But what should I actually do differently?” So this is the focus of this article. Understanding that quality is a distributed system is one thing, and actually designing your operating model around that understanding while in a real org, with real constraints and real people around you, with a tangible product that you are delivering, without that delivery slowing down for you while you try to figure out this new world around AI. And it seems that there are many leaders that are somewhere in this gap between these things – i.e. they understand things conceptually, but they haven’t built the operating model yet… So this article is hopefully a bit of a blueprint for closing that gap.

What an operating model actually is, and isn’t

Before we get into the model itself, let me describe what an operating model is, because the term gets loosely brandished about, and the looseness can cause misunderstandings and problems. An operating model is a system of decisions, roles, protocols, and feedback loops, that determines how something actually happens day to day. It’s not an org chart, or a process diagram, or a RACI matrix, (although these things might either feed in contextually, or be produced as an output relative to an operating model).

Most quality operating models that I’ve encountered in organisations answer four questions, whether or not they’ve been written down explicitly:

  1. What flows through the system?
  2. Who and what does the work?
  3. What does the system tell you?
  4. How does it get better over time?

The old “quality model” answered these questions in a particular way. Requirements flowed from a PO/PM to designers, developers, and testers in a certain sequence. A tester would embed into the agile team, and would test features and changes in defined ways. Developers would either sequence their testing by writing a test first then writing code to pass the test (with the test acting as a design artefact), or they’d sequence their testing by writing code, then writing the test to assert it was meeting expectations from the requirements from the user stories and acceptance criteria. For all this testing, a level of understanding of quality was presented along with pass and fail rates, and decisions were made on releasing features, etc. And retrospectives were there to supposedly improve the system, (when there was time to focus on actioning those ideas on how to improve it).

And I’m not picking apart this old system either – it made sense for the environment that it was designed for. The problem now though, is that the environment has changed, but it seems most quality operating models haven’t.

The AI-native version answers the same four questions differently, and that’s what the rest of this article maps out.

The five layers in practice

In article 1 of this series, I introduced a five-layer model for thinking about how quality flows through a modern engineering org. I wanted to build that out here into something more practical, because each layer has specific design decisions inside it that leaders actually need to make. It’s those design decisions that determine whether the operating model works (way more than the tools or the headcount!).

Layer 1: What flows in

Previously: Requests flow in from various sources (users, clients, C-Levels, the team, the product org, etc) which help form requirements (Epics, User Stories, etc) typically owned by the PO/PM, and handed over to developers and testers, roughly in some kind of sequence, with info and clarifications about “expected quality” flowing in that same direction.

New for AI-native: PRs, stories, behaviours, production signals, and AI-generated insights, past requirements, the info about the requests, etc – all build “context” which enters the system simultaneously, from multiple avenues, without any necessary handoffs, with AI agents able to read all this info directly.

The design decision leaders need to make: how does your team ensure AI agents have the right context to interpret what’s flowing in? Context quality is some form of new test plan. Who owns it? How is it maintained? What happens when a story is ambiguous or an acceptance criterion is missing? These aren’t tooling questions at all – they relate to the operating model itself, and they need to be answered to improve the output quality of the AI agents.

Layer 2: What connects it

Previously: people manually relay context between tools and other people, unless they are lucky to have some tools that offer 3rd party integrations or APIs between the tools they use. A tester reads the story, assesses its quality with regards to ambiguity, assumptions, risks and unknowns, etc, then perhaps supports in adding acceptance criteria, and in writing a tactical testing plan. SDETs also tactically plan, then design and write some automation tests. Developers plan who will review code, and what context they’ll need in order to do so effectively. Someone provisions the environment to be up and ready. There’s lots of handoffs, and every one of them is a potential point for there being a loss of information.

New for AI-native: AI agents reading across tools, interpreting intent, and generating scripted tests and insights. The connection between “what was intended” via any kind of explicit expectations or descriptions and “what gets tested” becomes more direct rather than mediated.

The design decision leaders need to make: Which connectors are you automating which need a human in the loop, and how do you know when to change that boundary? This is a living boundary too, that should move as trust in the system grows. It’s also worth noting that a team that has been using AI within their testing for 3+ months will have different answers to these questions than a team who has only started using AI in their testing last week.

Layer 3: What runs it

Previously: a testing phase with a defined start and end – be that once some feature code has been written and can be run by the tester embedded in the team, or even as the code is being written if a team has the privilege of having testers and developers able to pair and work together collaboratively. Testing is still typically owned by a tester though, and there is somewhat of a handoff into and out of that testing activity. Let me explain what I mean… Although the phrase “continuous testing” has been around for a long time, and has 2 distinct meanings in the testing world (either “continuously applying testing activities holistically throughout the SDLC”, i.e. testing ideas, requirement artefacts, designs, code, the product feature or change, the release processes, etc, or “repeatedly running a specific suite of automated scripted tests in a pipeline”), in both of these descriptions testing is a series of activities that have a set of distinct beginning and end criteria.

New for AI-native: continuous execution running in the background without a defined phase, but the more important shift is what execution now means. In the AI-native model, layer 3 contains three distinct modes running in parallel:

  • Human-led execution, the exploratory testing, risk investigation, and judgement calls that only people can make.
  • Deterministic automation, scripted tests and CI pipelines doing what they’ve always done reliably.
  • Agentic execution, AI agents actively exploring the application, generating and running tests, identifying defects, and adapting coverage without a human directing every step.

Testing is something that’s always occurring, and shouldn’t be thought of as something that happens just before release. And crucially, these three modes coexist rather than one replacing the others.

The design decision leaders need to make: what should each mode own, and where are the boundaries between them? Human judgement belongs in the exploratory and risk-investigation work that requires curiosity, creativity, and context. Deterministic automation belongs in the verified, high-confidence regression paths. Agentic execution belongs in the broad, continuous coverage work that would otherwise be impossible at the pace development actually moves. Getting that division right is the real design challenge in layer 3.

Layer 4: What it tells you

Previously: With scripted testing, you’d get information about pass and fail rates and metrics, coverage percentages, defect counts and other numbers that tell you about the correctness or incorrectness of the software or feature. With exploratory testing, you’d also get testing notes that shared qualitative information about what was investigated, how it was tested, and any discoveries related to problems or defects, any new risks. You’d also get a perspective of quality related to the scale of goodness in the experience of using the software, and also ideally a scale of anticipated value of the product.

New for AI-native: risk signals, coverage information (across all the shared context, not just against acceptance criteria points), flaky test detection, even potentially any compliance evidence. Information that tells you where the danger is sitting currently, which parts of the system are under-tested relative to their rate of change, and what the delta is between this week and last week.

The design decision leaders need to make: what does your intelligence layer need to surface for you to make good release decisions? Most leaders, if they’re being honest, are making release decisions based on limited lagging information, and maybe even on gut feel and relationship trust. What would it look like if you had a real, fuller picture of risk instead?

Layer 5: What comes back

Previously: a retrospective session (occasionally, when there was time). Improvements being fed back into the system about the quality of testing within the system (albeit based on the variable mindsets of what testing is by the people giving the feedback). Occasional updates to the “testing strategy” at a cadence of every 6-12 months.

New for AI-native: Continuous feedback into the PRs, stories, dashboards, and leadership decisions. The system learns from every cycle, agent run, and discovery that gets actioned or dismissed. The feedback loops are built into the workflow directly.

The design decision leaders need to make: how do you close the loop in a deliberate way? Who reviews what the system learned and decides what to do with it? If nobody owns the feedback layer, the system runs but  won’t improve. A system that runs but doesn’t improve will eventually stop being trusted for the exact reasons I mentioned in article 2 of this series.

Where BearQ sits in this model

To clarify off the bat: BearQ isn’t an operating model. No single tool is an operating model in the context of how I am describing it in this article. But BearQ yet again serves as a great example of where an agentic testing product has been designed to operate meaningfully across all five layers of the quality operating model, rather than just one or two layers.

Here’s how I see BearQ mapping across the layers.

Layer 1: BearQ is both a consumer and a contributor. It reads the context that you give it access to, and over time its application model starts feeding AI-generated insights back into the system, informing what gets prioritised in subsequent cycles. The intelligence it builds compounds with each feedback loop.

Layer 2: The Explorer and QA Lead agents handle the connective work, reading the application, interpreting intent from stories and PR comments, generating tests that attach back to their source without a human in the relay chain. This is still one of BearQ’s strongest capabilities.

Layer 3: BearQ’s agents are actively doing testing activities, exploring the application, generating and executing tests, identifying defects, and adapting coverage as the application changes. So it’s agentic execution and deterministic automation, sitting alongside human-led exploratory testing in the same layer. The three modes coexist. BearQ handles the agentic mode while connecting systems or reporting back.

This is also where SmartBear’s mission around “application integrity” starts to feel real as a continuous, measurable assurance that software works as intended, with the governance needed to operate at AI speed and scale… That’s a property that has to emerge from the execution layer.

Layer 4: The dashboards, risk clustering, coverage intelligence, and structured defect reports turn what the agents found into something people can read and act on. This is the intelligence output the system produces.

Layer 5: Because BearQ’s application model gets smarter with every cycle, it genuinely contributes to the feedback loop. Coverage adapts based on what was found. Risk prioritisation shifts as the system learns more about the application. That compounding intelligence is one of BearQ’s most distinctive characteristics, and it’s what separates it from tools that run and report without learning.

The human-in-the-loop controls across all of this… the focus areas, guardrails, work reviews, etc, mean teams can dial the autonomy level up or down depending on where they are in the trust-building journey I talked about in article 2. That’s a big design choice that reflects how operating models actually change incrementally (with evidence) in real organisations.

If you want to try out BearQ, they have a 7-day free trial available now.

The roles that change

There’s a cyclical nature between operating models and the roles within an organisation – the operating model is designed to be used by people in the roles within the organisation, and the roles are designed by the org in line with the operating models and structures that the org desires to work within. And the latter part seems to be the bit that many leaders underinvest in. (I suspect it’s because of the pressures and upward visibility of choosing a tool forcing the priority).

Quality Engineers are the role I want to talk about most carefully here, because I’ve seen this framed badly elsewhere and I understand the misconceptions surrounding this role, and I don’t want to feed into that by mistake. The shift for QEs is towards something much more strategically significant, and it shouldn’t be thought of as a diminishment of the role. In the AI-native operating model, Quality Engineers direct the AI agents have 5 areas of focus:

  1. Defining and refining the context (and the quality of the context too), and the constraints that shape what the agents use.
  2. Directing AI agents towards the right areas of the application.
  3. Conducting exploratory testing and risk investigation in the areas where human curiosity and judgement matter most.
  4. Reviewing agent output quality with genuine evaluation utilising their lateral and critical thinking skills.
  5. Interpreting risk signals from the intelligence layer, and improving the system’s understanding of the application over time.

The above can also be thought of as a “quality system designer” role within the org… with QEs bringing their expertise, judgement, and systems level thinking into play. These activities above need the skills that great QEs already have – it’s just a case of applying these skills at a higher level.

Developers become active contributors to the quality system’s intelligence, and that’s a more specific shift than it sounds. In the old model, the quality system didn’t really need much from developers beyond the code itself. In the AI-native model, what developers write around the code, the PR description, the acceptance criteria, the behaviour note left in a comment, that all becomes the context that shapes what the agents test and how well they test it. Vague acceptance critera produces shallow agent coverage. Precise criteria, with clear expected behaviours and named edge cases, produces something much more useful.

So this changes the discipline of writing stories and PRs. It’s a quality skill, not just a communication and design one. Developers who understand what an agent needs in order to do its job well will get faster, more specific feedback on their work. Developers who treat PR descriptions as an afterthought will find the agent’s coverage reflects that.

There’s also a shift in the developer and Quality Engineer relationship that I think is underrated. In the old model, that relationship often played out at the end of a cycle, in review or defect triage, with some friction baked into the handoff. In the AI-native model, Quality Engineers are setting context and constraints that shape what gets tested against the developer’s work from the start. The conversation moves earlier, it becomes more collaborative, and a lot of that end-of-cycle friction just stops existing, because there’s no handoff to trigger it.

Engineering Leaders face the most structural shift of the three roles, and I think it’s the one that gets talked about the least. Their quality responsibility was always distributed across embedded Quality Engineers and developers, but informally, relying on people to self-coordinate. The AI-native operating model makes that distribution visible in a way that can’t be managed informally anymore.

Specifically, I see the leadership role changing in three ways:

  1. The first is moving from reviewing quality artefacts to designing the system that produces them. Test plans, defect reports, coverage metrics, all of that used to land on the leader’s desk for review. In the AI-native model, the leader’s job is to design the system that generates those signals, not to inspect them after the fact. That means making the five-layer design decisions explicit across all their teams, not assuming each squad will figure it out independently. If Layer 2 is wired differently in every team, the intelligence in Layer 4 can’t be compared or used collectively at the leadership level.
  2. The second is moving from team-level quality thinking to system-level quality thinking. Individual teams can optimise their own testing without the organisation having any coherent view of where the risk is actually sitting. The Engineering Leader’s job is to set the protocols, the shared definitions of risk, the quality gates, the standards for what the intelligence layer needs to surface, so that distributed execution produces a coherent picture rather than five separate ones.
  3. The third is the one that requires the most personal change, and it’s about the quality conversations the leader is showing up to. If the sprint review is still asking “how many bugs did we find this sprint?”, the signal being sent is that the old model still applies. The questions that pull the system forward sound more like: “what is the risk picture across the product right now?”, “which areas of the application are the agents least confident about?”, “where is test debt accumulating relative to change rate?”, “are our Quality Engineers directing the agents well, or are they still doing work the agents should be handling?”

That last one matters most, because the Quality Engineer’s evolving role only actually happens if the leader is creating the conditions for it. If the Engineering Leader is still measuring Quality Engineers by defect counts and test case volume, the role won’t shift regardless of what tools are in the pipeline.

And one thing I want to be really plain about: none of these role shifts happen by announcing them. They happen by changing the system people work inside, which is the leader’s job to design.

How to actually start

This is the section I most want to get right, because “here’s a five-layer model” is only useful if there’s some kind of plan for how to move from where you are now toward that. And my strong advice is to treat this as a sequence, not a simultaneous rebuild, because trying to redesign all five layers at once is one of the faster ways to exhaust your team’s goodwill and produce very little that’s visible or trustworthy.

The place I’d start is Layer 4, and that might feel counterintuitive. Before changing anything about how testing happens, change what you measure and what you look at. Swap defect counts for risk signals, and swap coverage percentages for coverage trends. The reason to start here is that it builds the leadership habit of reading the system rather than counting its outputs, and it also builds the case for everything else, because once you start seeing the gaps in your current intelligence picture, the need for a proper connective layer (Layer 2) stops feeling theoretical and starts feeling urgent.

Once you have some traction on Layer 4, Layer 2 is where I’d go next. Start with a single, bounded connection, PR-driven test generation is usually the most accessible entry point because the feedback loop is tight and the value shows up quickly. The temptation is to connect everything, but connecting one thing well and building trust in that connection is worth far more than half-connecting five things and confusing everyone.

Layer 5 is the one most teams skip, and it’s the one that determines whether the system gets smarter over time or just keeps running at the same level. Someone needs to own the question of “what did the system learn this cycle, and what are we going to do about it?” That doesn’t have to be a formal process, it can just be 15 minutes in a retro or a quick async note, but it needs a home and it needs an owner. Without it, the feedback loop doesn’t close, and a system that doesn’t learn will gradually stop being trusted for exactly the reasons I described in article 2.

Layers 1 and 3 are the ones I’d leave until last, partly because they’re often already working to some reasonable degree (requirements are flowing in through some tool, tests are running in some pipeline), and partly because adding sophistication to Layer 1 before you have a Layer 2 that can interpret it is just generating noise. Get the connective and intelligence layers stable first, then go back and optimise what flows in and how execution is structured.

And throughout all of this, run the role shifts in parallel with the system changes. A system designed for an AI-native quality model, but staffed and measured against the old one, will consistently underperform and give ammunition to everyone who was sceptical about it in the first place.


The operating model is the strategy

I want to close with something I genuinely believe doesn’t get said enough in conversations about AI and quality.

The organisations that lead on software quality over the next few years probably won’t be the ones with the most sophisticated AI tools. They’ll be the ones with the most deliberately designed quality systems. Clear layers. Clear roles. Clear feedback loops. Leaders who understand what the system is telling them and make good decisions off the back of it.

BearQ is a genuinely strong example of what the connective and intelligence layers can look like when a tool has been built with that kind of intentionality. But the operating model around it, the design decisions in each layer, the role evolution, the sequencing, the feedback loops, that’s all still something leaders have to design themselves. No tool does that for you.

The quality operating model is central to how engineering organisations deliver software in the AI-native era, and it shouldn’t be thought of as a supporting function anymore. Leaders who design it deliberately will move faster and with more confidence than those who let it emerge by accident, because in complex systems, what emerges by accident is usually entropy.

There’s one final piece that makes all of this defensible at scale though, especially if you’re working in a regulated environment, or in an organisation where trust in AI systems is still being earned. That’s governance. And governance is something that enables a quality operating model to run at speed, and it shouldn’t be thought of as the thing that slows everything down. That’s where the final article in this series is headed.

Please leave a comment!