Trusting AI Series | Blog 3

Trusting AI for Business Outcomes

Matt Scarborough CEO and Co-Founder
August 13, 2026 11 min read

The inspiration for this blog series came from discussions at the CDIA Connect conference in May of this year, during which the question “Can I trust AI?” was a common topic.

This post is longer than the first two because it tackles the biggest conceptual shift: AI expands the range of business processes that can be improved through automation, but that expansion also requires a different way of defining success, measuring quality, and managing residual risk compared to prior eras of automation.

One caveat before going further: I have learned quite a bit about AI tools in recent years and can describe what we are building with confidence, but the field is still moving quickly. I expect the principles in this post to remain useful, but the specific practices, tools, and controls continue to evolve.

As described in the first blog of this series, it really struck me that meaningfully answering the “Can I trust AI?” question starts with completing it by adding two key parts:

  • To do what?
  • How well?

I’ve spent most of my career at the intersection of business functions and technology solutions for consumer lending products, and the importance of the questions above is not new:

  • Successfully improving business using technology has always required a clear understanding of the desired business outcome and related requirements.
  • AI does not change that.
  • It does, however, change the way we answer those questions in very significant ways.

The rest of this post follows that logic in three steps.

  1. First, I’ll explain what has and has not changed about defining intended business outcomes.
  2. Second, I’ll explain why AI requires a different way of thinking about quality, especially where compliance and audit teams already have expectations for how a process should be evaluated.
  3. Finally, I’ll connect those ideas to the question of how deeply an organization needs to review an AI-enabled solution before relying on it.

What Changed

Let’s explore what changed, and what didn’t, before we get into how to adapt to the changes.

To Do What?

When I started my career in the 90's, anything we asked technology to do had to be very specifically defined. The methods of defining it evolved over time, from waterfall requirements to Agile frameworks or some hybrid, but the outcome was the same: computers required specific instructions, so the business had to eventually define exactly what should happen.

Getting those requirements right was critical and hard to do. Failing to do so was the root cause of many disappointments and real tension between the business and IT sides of organizations.

The move to Agile changed both the method and the emphasis. The focus shifted from up front development of detailed technical requirements to using iterations of working software to help discover what the functional requirements truly were.

This eliminated the need for lengthy requirements at the beginning, but doing it right still required clarity about the desired outcomes: defining the business results to be accomplished and how they'd be measured, then iteratively developing software to achieve part of those results.

With generative AI, the principles that drove the shift toward Agile apply more than ever. Agile works well when building a prototype and then changing it is fast compared to debating detailed requirements up front, and generative AI makes that kind of rapid prototyping faster still.

Doing it well, though, still requires clarity about the desired business outcomes. What really changes is the breadth of potential use cases. We're no longer limited to those that can be defined in completely deterministic code. This opens the door to automating processes that were too complicated to be realistic candidates before.

How Well?

Because the only processes that could be automated in the past were ones that could be precisely defined, the final product was expected to be executed with very high reliability. Targets were often six-sigma or better: three defects per million or getting it right 99.9997% of the time. Everything that couldn't be defined so carefully remained a manual process.

Now that we can automate processes that would have otherwise stayed manual, we need to think about quality differently. If a manual process is done right 80% of the time, an AI-enabled solution getting it right 95% of the time is a huge improvement, even while falling short of the reliability we expect from traditional automation.

It's therefore critical that the "How well?" measurements applied to an AI solution be the same ones used to judge the manual process it replaces. In some cases, AI even enables quality metrics that didn't exist before. When it does, apply the new metric to both the AI solution and the manual process.

For example, few lenders today comprehensively measure how well their agents read the attachments borrowers submit in credit bureau disputes. AI may let a lender identify that information with far greater consistency than a manual process and measure the difference. Relevant accuracy of metrics will therefore mix percentages, which we're used to doing for technology, and percentiles, which we historically haven't used much.

Imagine a process where the AI solution is right 87% of the time, but that performance sits in the 97th percentile compared to manual work. Being right 87% of the time would have been unacceptable in prior use cases, but if the process is challenging enough that this beats 97% of manual results, it's a major upgrade.

Knowing both numbers matters: leadership can be confident that average quality will improve, while the combination also reveals opportunities to route certain characteristics for manual review and to keep pushing that 87% higher.

Illustrative Example

Let's now briefly illustrate the connection between what we're asking technology to do and how well it must do it. Imagine you're planning a vacation, and consider how different your answer would be to each of the following questions:

A Simple Risk Spectrum

“Can I trust AI to…”

  • …help me brainstorm potential locations?
  • …summarize a handful of articles about potential destinations well enough for me to decide whether any are worth reading closely?
  • …iteratively plan a general itinerary based on the locations I choose, the expected weather, and our likes and dislikes?
  • …make specific bookings on my behalf for flights, ground transportation, hotels, and restaurants after presenting options and receiving my confirmation?
  • …fully execute certain bookings without confirming with me, based on guidelines I've set?
  • …drive me and my family at relatively low speeds within a town we visit?
  • …drive our car on a mountain road in an ice storm, without me even having a steering wheel?

The answers vary greatly across those questions, as does the implied answer to "How well?", since the consequences of getting things wrong changes so significantly.

Said another way, the inherent risk changes radically across the spectrum above, so our strength of controls would need to change just as radically to result in acceptable residual risk.

Applying This to Our Own Products

The same is true when evaluating solutions in our professional world. I'll use two of our AI-enabled products as examples.

Our AI Research Assistant , a premium feature within our DQS Furnishing module, helps with what is generally an ad-hoc investigative or querying process. The relevant high-level question is:

"Can I trust the AI Research Assistant to provide factually accurate answers and observations well enough to help a skilled user investigate faster and more comprehensively than they otherwise would?"

Its business outcomes should be evaluated in that light. It would not be appropriate for creating executive-level reports without significant further review, but neither would the output of an ad-hoc query by a human analyst.

Credit Bureau Disputes

AI Resolution Engine

Our AI Resolution Engine for credit bureau disputes assists human agents in addressing those disputes. Here the relevant question is:

"Can I trust the AI Resolution Engine to provide accurate information, instructions, and recommendations to our agents well enough to materially increase the quality and efficiency of our dispute investigations?"

In this case, organizations already have metrics and QA/QC processes in place to evaluate dispute resolution, so the key questions are whether the solution improves the metrics we use today, and whether it lets us measure quality in ways we couldn't before, across both automated and manual processes.

This is where the connection with compliance becomes especially helpful. Credit bureau disputes receive significant attention from regulators, litigators, and therefore internal compliance and audit functions, so the definitions of what must be done, and how well, are better established than for most processes.

Somewhat counterintuitively, this makes a highly regulated process like FCRA dispute investigation a strong candidate for AI-enabled automation: the obligations, quality expectations, and evaluation concepts are already clearer than in many less regulated processes.

The productivity gains matter, but the potential for greater quality and consistency means compliance and audit stakeholders can be strong advocates for these efforts. For a group already accustomed to defining a "reasonable investigation," the measurement concepts described earlier in this blog are very familiar.

Together, these examples help define what we're trying to trust AI to do, how well it needs to do it, and why compliance and audit stakeholders can be important partners in that evaluation.

The Next Question: How deeply do you need to investigate before responsibly relying on an AI-enabled solution?

The depth to which an organization must review the performance of any technology, including AI tools, depends on two broad factors: the purpose of the solution and the inherent risk of the activities tied to that purpose, and the organization's role in creating and/or using the solution.

I believe confusion about the appropriate depth of analysis is one of the main challenges organizations face as they adopt AI. Evaluating too shallowly can leave significant unknown risks and lead to failures, while evaluating too deeply can create confusion and waste time.

Let's illustrate this with an analogy: buying a rain jacket. The right depth of analysis depends both on your role and on how much is riding on the outcome.

Imagine a new style from an established company that you plan to wear every day. Because it's new, there are no reviews yet. The questions that you would likely ask as a consumer are very different than the questions that should be asked by other people involved in the process of creating that jacket and making it available to you.

What would you likely ask, and what questions would the others need to think about before you ever see it?

Role
Example questions they'd ask
Role Consumer
  • What claims does the company make? e.g., "Is this both waterproof and breathable?"
  • Does it use components I already trust? e.g., "I know I trust Gore-Tex."
  • What are the reviews on related products from this company?
Note

We first imagined the jacket for everyday use, where the inherent risk is simply being uncomfortable if it performs poorly. If instead you were taking it on an expedition where gear failure could be fatal, you'd investigate more deeply than a typical consumer, borrowing some questions from an adjacent role, but not the entire stack.

Role Buyer for a large retailer
  • How was the garment assembled, and what documentation supports it? i.e.: "Is this seam-sealed, and how?"
  • Can the manufacturer deliver on time and on budget?
Role Engineer at the manufacturer
  • Can we improve the breathability of the moisture barrier by adjusting the temperature, speed, or chemical inputs in the manufacturing process?
Role Chemistry researcher
  • What compound characteristics create the molecular structure behind the moisture barrier — gaps big enough to let water vapor through, but too small for liquid water?
  • Can we design molecules that perform better, or perform as well at lower cost or environmental impact?

The key is this: Ask the most detailed questions about the things you know the most about. If you know the business outcomes that must occur and how to measure them, then be demanding and specific on those points. If the solution meets those metrics, then great. If not, then it is not good enough.

The questions at the deeper layers will need to be addressed by the chain of people delivering the solution to you. In the example above, that “solution” is a rain jacket. In the case of AI, this means focusing on what you know and not letting questions about the depths of the model be a distraction from the key questions about how it works, or not, for your use case.

A consumer who tried to answer all of these questions before purchasing would likely never buy a rain jacket at all (and would just get wet).

The Practical Takeaway

The Bottom Line

The practical takeaway is that trust in AI needs to be narrowly and precisely defined to have meaningful answers, and for business and compliance professionals, the focus belongs on business outcomes. Which model is being used matters far less than how well the overall solution works, measured against existing quality metrics.

I'll pick this up in the next blog, focusing on what it takes to roll out real AI-enabled solutions responsibly. I'll expand on these themes using our AI Research Assistant and AI Resolution Engine for credit bureau disputes, since the similarities and differences between them show how the right business outcome, quality measures, control expectations, and stakeholder roles vary by use case.

My commitment here isn't perfection. It's transparency, disciplined measurement, and a willingness to keep improving as the tools, the controls, and the industry's understanding continue to evolve. As always, please let me know any feedback you have.

Author

About the Author

Matt Scarborough

CEO and Co-Founder

Matt Scarborough is the CEO and Co-Founder of Bridgeforce Data Solutions, a RegTech SaaS company helping financial institutions, credit unions, and fintech lenders improve credit reporting accuracy and automate dispute resolution. Since co-founding the company in 2016, Matt has helped lead the development of industry-leading solutions built around the complexities of Metro 2® data, credit bureau disputes, and regulatory compliance.

More recently, he has guided the team on the thoughtful development of new optional AI capabilities designed to help clients address these challenges more effectively.

Matt is an active speaker on AI applications in credit reporting and compliance, including sessions connected to the AI-Native Banking & FinTech Conference, ACU’s Regulatory Compliance School, and CDIA Connect. His talks focus on practical AI implementation, iterative deployment, and regulatory alignment in consumer credit and dispute management.

Related speaking-session articles

Complaint monitoring and regulatory change New AI findings from ACU FCRA, dispute abuse, and AI insights

Frequently Asked Questions

Does an AI solution need to be perfect to be trustworthy?
No. AI should be measured against the process it's improving, not against six-sigma automation standards. An AI-enabled solution that's right 95% of the time can be a major upgrade over a manual process that's right 80% of the time, even though 95% would fall short of the reliability expected from traditional automation.
Why are highly regulated processes like FCRA disputes good candidates for AI?
Because the obligations, quality expectations, and evaluation standards are already clearly defined by regulators, litigators, and internal compliance and audit teams. That clarity gives AI-enabled automation a well-established bar to be measured against, which is why compliance and audit stakeholders can be strong advocates for these efforts rather than obstacles to them.
How deep should you investigate before relying on an AI-enabled solution?
It depends on the use case. For an ad-hoc tool like the AI Research Assistant, a skilled analyst checking the output against their own judgment is usually enough. For a tool like the AI Resolution Engine, where outputs inform live dispute decisions, that review needs to run deeper, through the same QA/QC and compliance processes already used for human-led investigations.