Auditing an AI-enabled data solution requires reproducible evidence

MATOS AI turns audits of AI-enabled data solutions into reproducible evidence covering architecture, data, models, controls, and operations.

This article is also available in Spanish.
Auditing an AI-enabled data solution requires reproducible evidence

What it is

The architecture committee approved a solution that classified master records and corrected anomalies with model assistance, while its dashboard displayed availability, latency, and an aggregate accuracy metric. During the review nobody could identify which evaluation dataset version supported approval, which transformation produced the input features, or which policy authorized automatic corrections, so the solution was responding and its components were deployed even though the available evidence could not reconstruct why a decision had been accepted.

An AI enabled data solution can only be considered audited when a competent third party can reconstruct which version operated, which data it used, which objective governed it, which result it produced, which controls intervened, and which mechanism would have contained a failure. The audit must examine the full system because observed behavior depends on the data pipeline, preparation rules, model, prompts, retriever, tools, access policies, human interface, and third party services.

Two scopes should be separated in the audit mandate because auditing a solution that uses AI means evaluating a platform whose operation depends on predictive, generative, or decision models, or that applies AI to integration, quality, classification, monitoring, or metadata management. Using AI to support an audit means employing a model to review documentation, classify findings, select samples, generate queries, or detect patterns, so the first scope examines the production system and the second introduces an additional tool that also requires controls, version records, and independent validation.

The proposed framework is called the Technical and Operational Audit Framework for AI Enabled Data Solutions or MATOS AI and it does not replace NIST AI RMF, ISO IEC 42001, ISO IEC 42005, ISO IEC 42006, ISO IEC 23894, or applicable auditing standards. Its purpose is to turn those references into a technical procedure that connects control objectives, artifacts, tests, results, residual risks, and acceptance decisions.

MATOS AI distinguishes the existence of a control from its effectiveness because a document may state that a model will be monitored while the audit must demonstrate that the metric is calculated over the correct population, that the threshold was approved, that the alert reaches an accountable person, and that a real or simulated incident produces a verifiable response. The review moves from documentary presence to design adequacy, implementation, effectiveness over a period, and reproducibility.

The audit boundary must include assets that alter the outcome even when they sit outside the main repository, from training data, labels, transformations, features, artifacts, and rules in a predictive model to foundation model, prompt, index, retrieved documents, filters, and evaluators in a generative application. An agent also includes tools, credentials, memory, action sequence, approvals, and compensating operations, so an audit limited to the model file leaves several likely failure causes outside its scope.

What it adds

The first benefit is replacing declarative trust with test based assurance, using the GAO framework that organizes AI accountability around governance, data, performance, and monitoring and assigns questions and procedures to auditors and assessors. NIST AI RMF adds a continuous view through Govern, Map, Measure, and Manage, where deployment, maintenance, or retirement decisions depend on the stated purpose and measured risk throughout the lifecycle.

The audit also prevents an aggregate metric from hiding material failures because a classifier may retain acceptable overall accuracy while error increases for a rare category and a generative system may answer fluently while citing expired documents. A forecasting model may preserve its average error while degrading a region, retail chain, or product family, so a useful test preserves segmentation, period, data version, confidence intervals where applicable, and the relationship with the decision changed by the system.

Australia's Robodebt case provides a documented reference for automation, data, and institutional control because the Royal Commission recommended stronger documentation and legality reviews for data exchanges, publication of business rules and algorithms for expert scrutiny, and a capacity to monitor and audit automated decisions across technical aspects, fairness, bias, and usability. The consequence for an AI data audit is concrete because isolated accuracy testing does not cover legality, appeal, business rules, or effects on people.

MATOS AI provides a common structure for different solutions without imposing identical tests on all of them because a demand prediction that only informs human review requires controls different from an automated decision that blocks a payment. An assistant that proposes column mappings can tolerate manual rejection and rework, while an agent that deletes records requires strong authorization, idempotency, batch limits, and rollback, and materiality is determined by decision severity, autonomy, volume, data sensitivity, reversibility, and propagation speed.

A design audit determines whether proposed controls cover the anticipated risk and a preproduction audit verifies that the released package matches the evaluated one. An operating effectiveness review observes a period and confirms that controls worked across real data, changes, alerts, and incidents, while a change or incident triggered review concentrates testing on the modified area and its dependencies when data, provider, foundation model, prompts, rules, or permissions change.

Using AI to support an audit can increase coverage across large inventories, extensive logs, and dispersed documentation because a model can suggest ownerless assets, group recurring errors, or propose high risk samples even though its output acts as an audit lead rather than sufficient evidence. The auditor must preserve the model, version, prompt, supplied data, enabled tools, and subsequent validation because a nonreproducible response or instructions embedded in reviewed documents can create false findings or omit relevant exceptions.

There is an operational limit because a deep audit consumes specialist time, evidence storage, instrumentation, security testing, and business owner participation, while applying the same depth to an internal prototype without material decisions may consume resources without reducing an equivalent risk. Scope should remain proportional, although reduced testing must be justified in writing and linked to conditions that require broader review when autonomy, affected population, data exposure, or operational dependency increases.

How to implement

MATOS AI starts with a mandate that defines the object, purpose, period, population, affected decisions, actors, providers, and materiality criteria, while the inventory must associate every component with an owner, a version, and obtainable evidence. Technical review covers eight domains, and each domain must end with an observable condition to approve, restrict, or stop.

Audit domain Control objective Minimum evidence Representative test Stop condition
Scope and accountability Intended use, boundaries, and owners are approved Mandate, RACI, inventory, risk classification, and acceptance decision Compare declared use with real calls, users, and decisions The system performs an unapproved use or lacks an accountable owner
Architecture and dependencies Every component that changes the outcome is identified Repositories, integration documentation, versions, contracts, and providers Reconstruct the route from input to decision and response A production dependency has no version, owner, or exit mechanism
Data and lineage Data is legitimate, adequate, representative, and traceable Sources, permissions, contracts, hashes, transformations, quality, and retention Reproduce a sample from source to features or retrieved context Origin, transformation, or authorization of material data cannot be demonstrated
Model and lifecycle The deployed artifact matches the evaluated artifact and can be reproduced Code, parameters, image, digest, registry, dataset, and evaluation report Rebuild the artifact and compare metrics within approved tolerances The production version does not match the evaluated package
Performance and bias Metrics represent the objective and material groups Segment metrics, calibration, errors, human evaluations, and thresholds Repeat evaluation across segments, periods, and adversarial cases A critical group or process exceeds risk tolerance without approved treatment
Security and privacy Data, models, and tools resist misuse and excessive access IAM, secrets, tests, logs, filtering, encryption, and threat analysis Run negative tests for injection, exfiltration, and tool abuse The system permits access, action, or disclosure outside the authorized purpose
Human control and decision Human intervention has enough authority, information, and time Decision matrix, review queues, reasons, overrides, and appeals Sample decisions and verify that review can change the outcome Human approval is nominal or cannot stop a material action
Operations, cost, and retirement The system can be monitored, contained, rolled back, and retired SLOs, alerts, runbooks, incidents, costs, canary, rollback, and retirement plan Simulate degradation, provider loss, revocation, and rollback There is no proven way to limit damage or return to safe operation

Evidence should preserve a chain linking a production outcome with its components. An audit manifest can materialize that relationship and prevent the review from depending on screenshots or mutable names.

audit_case_id: AIA-2026-014
system_id: master-data-classifier
assessment_period: 2026-06-01_2026-06-30
intended_use: propose_product_classification
risk_tier: medium
release:
  git_commit: 8af31c2
  image_digest: sha256_7a93d1
  model_uri: models_classifier_17
  prompt_hash: sha256_39b21e
  policy_version: policy_12
inputs:
  training_dataset_hash: sha256_b811d4
  evaluation_dataset_hash: sha256_2942ac
  feature_contract: product_features_v6
controls:
  human_approval: required
  automatic_write: false
  rollback_runbook: RB-CLASS-04
evidence:
  evaluation_report: eval_2026_06_28.json
  bias_report: slices_2026_06_28.json
  lineage_run_id: ol_01J2M8P
  approval_ticket: GOV-1842

The manifest does not demonstrate that the control worked by itself because every identifier must resolve to an immutable artifact that is accessible to the auditor and protected from later modification. Aliases such as production or champion simplify operations, yet evidence must preserve the concrete version referenced during the audited period because a silent change can make a reproduction use a model different from the one that produced the original decision.

The reproducibility test selects an execution or sample of decisions and rebuilds the data, transformations, configuration, artifact, and policy, comparing exact output for deterministic models when the environment permits. Stochastic components require controlled parameters, preservation of the original response, and repetition of an evaluation suite over a versioned set because demanding identical text may be the wrong criterion, while tolerance must be approved before the test and linked to the business decision.

Metrics vary by AI type because a predictive model needs error, discrimination, calibration, and segment performance together with a strategy for delayed labels, while a generative application needs retrieval quality, documentary grounding, source freshness, safe refusal, information exposure, and behavior under malicious instructions. An agent adds action legitimacy, tool trajectory, confirmation, idempotency, limits, and rollback success, and a data quality component should measure suggested rule precision, false positives, rejected records, human intervention, and postchange consistency.

Operating effectiveness requires observing the control over a period rather than during a demonstration, connecting alerts to tickets, tickets to decisions, decisions to accountable owners, and corrections to a deployed version. An alert left open, a threshold changed without approval, or an incident closed without validation evidence shows that the control exists in the interface but is not sustaining operations.

AI used inside the audit must operate on a separate track where it can read evidence with restricted access, propose queries, detect anomalies, and classify documents, but it should not modify sources or issue the final opinion. Its configuration enters the audit file and its results are compared against deterministic queries, manual sampling, or expert review, while reviewed documents are treated as untrusted content because they may contain instructions capable of manipulating an assistant with tools or broad access.

A next day verification can select ten production decisions and reconstruct for each one the request identifier, model or prompt version, data snapshot or hash, applied policy, response, human intervention, and subsequent outcome. If the team can only recover technical logs or an approximate version, the solution lacks sufficient evidence for an operating effectiveness conclusion and the first break usually appears between the decision and the artifact when datasets are overwritten, aliases move, or a provider is updated without preserving the previous configuration.

The audit output must separate the finding, residual risk, and decision because a missing control in an informative feature may produce a planned correction while the same absence in an automated decision about people or money may require immediate restriction. Acceptance needs an authorized owner, review date, compensating controls, and follow up evidence, and a finding is only closed when a repeated test confirms that the control works under the conditions that produced the observation.

How it affects ROI and EBITDA

The effect on EBITDA appears in costs that are often dispersed across engineering, operations, security, legal, and support because a reproducible audit reduces manual investigation during incidents, avoids repeating validations that already have reusable evidence, identifies components without owners, and restricts changes that could produce rework, downtime, or data exposure. It also enables provider assessment through comparable parameters because access to versions, data, logs, tests, and exit mechanisms is required before dependency grows.

Return on investment improves when one evidence package supports architecture, risk, security, compliance, and change approval, allowing the organization to reuse common controls without assuming that an institutional certification proves the behavior of every solution. The benefit depends on automating evidence collection, hashes, lineage, and tests within the pipeline because an audit assembled manually at the end of every release increases lead time and encourages incomplete files.

The framework introduces costs that must be budgeted through immutable records, evaluation datasets, expert review, segment monitoring, adversarial exercises, and artifact retention, while a limited review may be adequate for low materiality systems. For irreversible decisions, sensitive data, or autonomous models, reducing evidence shifts expenditure toward incidents, disputes, and later reconstruction when the organization no longer controls every variable.

The decisive human capability combines audit judgment with data architecture, model risk, security, and process knowledge because no isolated specialty can determine whether a correct metric serves the wrong objective, whether legally available data is appropriate for training, or whether a technically plausible explanation allows a decision to be challenged. Independence also requires that the person who built a control is not the only person declaring its effectiveness.

The review reaches its destination when the team can return to the decision presented to the committee and reconstruct it without tacit memory, even if the dashboard continues to display availability and accuracy. Approval must rest on a file that identifies the data, artifact, policy, accountable person, and test that would have stopped the solution, turning a statement of trust into evidence another person can examine and repeat.

MATOS AI technical guide

The proposed guide, its supporting documents, and this framework are authored by Luis José Raigoso Valcárcel. All rights reserved.

Recommended resources

References

National Institute of Standards and Technology. (2023, January 26). Artificial Intelligence Risk Management Framework AI RMF 1.0. NIST. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-ai-rmf-10
U.S. Government Accountability Office. (2021, June 30). Artificial Intelligence - An Accountability Framework for Federal Agencies and Other Entities. GAO. https://www.gao.gov/products/gao-21-519sp
Royal Commission into the Robodebt Scheme. (2023, July 7). Report. Australian Government. https://robodebt.royalcommission.gov.au/publications/report