AI SYSTEMS DON't STOP CHANGING
AI Testing and Auditing Services
Many AI systems never stop changing, as their design is probabilistic. Something that doesn't operate on IF-THEN-ELSE and isn't locked in with temperature=0* will produce varying results each time. In a data pool where the data input itself doesn't follow a predefined structure, that is a necessary requirement. That's how patterns can be found, human language can be interpreted and expressed.
This makes the evaluation of AI systems, particularly those that operate on deep learning - which includes all Generative AI applications - particularly difficult. It isn't about gating a release and then keeping it stable until the next change, it is about evaluating it permanently, to understand if it still performs as expected and delivers the results that are desired.
At 9senses, we separate two disciplines: continuous testing and independent auditing.
Testing - identifying the "what"
Testing evaluates if a system expresses the desired behavior and ensure the right action if it doesn't.
For AI this begins as part of the use case definition, dominates through all the stage gates and never ends after release.
And because AI systems are designed to act faster than humans, on many more parameters, the "human in the loop" is not a real option for every decision AI makes.
In a nutshell, AI testing provides the ongoing evidence that the AI still performs as intended.
Auditing - exploring the "why"
Auditing provides an independent view and goes one level deeper. It expands the use cases beyond the usual test cases, by deliberately identifying their nature and - for example - not their exact wording.
Most importantly, it identifies the root causes of observed behavior and looks for remedies. In GenAI systems that can be the entire search architecture, the prompting, or the harnesses and controls.
This is why we created the 9senses GenAI Audit Framework, consisting of a quickly administrable Level 1 audit and a deeper Level 2.
Test continuously. Audit independently.
AI assurance is not one gate before go-live. Testing provides continuous evidence. Auditing provides independent assurance.
For business-critical AI, you usually need both.