{"id":226905,"date":"2026-07-21T09:48:30","date_gmt":"2026-07-21T09:48:30","guid":{"rendered":"https:\/\/www.9senses.ai\/?page_id=226905"},"modified":"2026-08-27T08:16:50","modified_gmt":"2026-08-27T08:16:50","slug":"testing","status":"publish","type":"page","link":"https:\/\/www.9senses.ai\/de\/testing\/","title":{"rendered":"AI Testing"},"content":{"rendered":"\n<div class=\"et_pb_section_0 et_pb_section et_section_regular et_block_section\">\n<div class=\"et_pb_row_0 et_pb_row et_block_row ns-hdr\">\n<div class=\"et_pb_column_0 et_pb_column et_pb_column_4_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_text_0 et_pb_text et_pb_bg_layout_light et_pb_module et_flex_module ns-eyebrow\"><div class=\"et_pb_text_inner\"><p>AI testing does not end at go-live<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_row_1 et_pb_row et_flex_row ns-hdr\">\n<div class=\"et_pb_column_1 et_pb_column et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_post_title_0 et_pb_post_title et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_title_container\"><h1 class=\"entry-title\">AI Testing<\/h1><\/div><\/div>\n<div class=\"et_pb_text_1 et_pb_text et_pb_bg_layout_light et_pb_module et_flex_module\"><div class=\"et_pb_text_inner\"><p>AI systems cannot be tested like conventional software and signed off once.<\/p>\n<p>Their outputs are probabilistic, their environment changes, and acceptable performance depends on the use case and its risks.<\/p>\n<p>Testing therefore has to establish what acceptable performance means, test it independently, and keep validating it throughout operation.<\/p>\n<p>Our approach is informed by established AI testing and risk-management practice, including ISO\/IEC TS 42119-2:2025, ISO\/IEC 25059, ISO\/IEC 23894, and the NIST AI Risk Management Framework.<\/p>\n<\/div><\/div>\n<\/div>\n<div class=\"et_pb_column_2 et_pb_column et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_code_0 et_pb_code et_pb_module\"><div class=\"et_pb_code_inner\"><div class=\"ns-vmap-editor-note\" style=\"position:relative;width:100%;overflow:hidden;aspect-ratio:5\/4;max-height:230px;\"><div style=\"position:absolute;inset:0;display:flex;align-items:center;justify-content:center;pointer-events:none\"><div style=\"padding:10px 12px;border:1px dashed rgba(170,180,205,.55);border-radius:6px;color:#8a93a6;font:600 12px\/1.4 sans-serif;background:rgba(8,13,25,.46)\"><div style=\"margin-bottom:6px\">Vector map \u2014 small<\/div><code style=\"display:inline-block;padding:3px 6px;border-radius:4px;background:rgba(127,140,170,.12);color:inherit\"><div class=\"ns-vmap-editor-note\" style=\"position:relative;width:100%;overflow:hidden;aspect-ratio:5\/4;max-height:230px;\"><div style=\"position:absolute;inset:0;display:flex;align-items:center;justify-content:center;pointer-events:none\"><div style=\"padding:10px 12px;border:1px dashed rgba(170,180,205,.55);border-radius:6px;color:#8a93a6;font:600 12px\/1.4 sans-serif;background:rgba(8,13,25,.46)\"><div style=\"margin-bottom:6px\">Vector map \u2014 small<\/div><code style=\"display:inline-block;padding:3px 6px;border-radius:4px;background:rgba(127,140,170,.12);color:inherit\">[ninesenses_vectormap #3]<\/code><div style=\"margin-top:6px;font-weight:400\">Max height: 230px. Preview on the live page.<\/div><\/div><\/div><\/div><\/code><div style=\"margin-top:6px;font-weight:400\">Max height: 230px. Preview on the live page.<\/div><\/div><\/div><\/div><\/div><\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_section_1 et_pb_section et_section_regular et_block_section ns-block ns-tabset\" id=\"principles\" style=\"max-width:1080px;margin-left:auto;margin-right:auto;border-radius:0;hyphens:auto;-webkit-hyphens:auto;-ms-hyphens:auto\" lang=\"en\">\n<div class=\"et_pb_row_2 et_pb_row et_block_row\">\n<div class=\"et_pb_column_3 et_pb_column et_pb_column_4_4 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_text_2 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><h2 style=\"display: inline; margin: 0; padding: 0;\">Four principles of AI testing<\/h2>\n<\/div><\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_row_3 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_4 et_pb_column et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide hovergroup\" id=\"trigger-define\">\n<div class=\"et_pb_blurb_0 et_pb_blurb et_pb_bg_layout_light et_pb_blurb_position_top et_pb_module et_flex_module\"><div class=\"et_pb_blurb_content et_flex_module\"><div class=\"et_pb_blurb_container\"><h3 class=\"et_pb_module_header\">Define success before you build<\/h3><div class=\"et_pb_blurb_description\"><p>Testing starts with the use case, not the finished system. Define what the AI needs to achieve, how reliably it has to achieve it, and which failures are unacceptable.<\/p>\n<\/div><\/div><\/div><\/div>\n<div class=\"et_pb_icon_0 et_pb_icon et_pb_module et_flex_module\"><span class=\"et_pb_icon_wrap\"><span class=\"et-pb-icon\">3<\/span><\/span><\/div>\n<\/div>\n<div class=\"et_pb_column_5 et_pb_column et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide hovergroup\" id=\"trigger-evidence\">\n<div class=\"et_pb_blurb_1 et_pb_blurb et_pb_bg_layout_light et_pb_blurb_position_top et_pb_module et_flex_module\"><div class=\"et_pb_blurb_content et_flex_module\"><div class=\"et_pb_blurb_container\"><h3 class=\"et_pb_module_header\">Build evidence, not demos<\/h3><div class=\"et_pb_blurb_description\"><p>A few good outputs show that an AI can work. They do not show that it works reliably. A proper test architecture separates development from independent validation and deliberately challenges the system beyond the cases it was built around.<\/p>\n<\/div><\/div><\/div><\/div>\n<div class=\"et_pb_icon_1 et_pb_icon et_pb_module et_flex_module\"><span class=\"et_pb_icon_wrap\"><span class=\"et-pb-icon\">3<\/span><\/span><\/div>\n<\/div>\n<div class=\"et_pb_column_6 et_pb_column et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide hovergroup\" id=\"trigger-baseline\">\n<div class=\"et_pb_blurb_2 et_pb_blurb et_pb_bg_layout_light et_pb_blurb_position_top et_pb_module et_flex_module\"><div class=\"et_pb_blurb_content et_flex_module\"><div class=\"et_pb_blurb_container\"><h3 class=\"et_pb_module_header\">Release is a baseline<\/h3><div class=\"et_pb_blurb_description\"><p>Go-live establishes how the system performed at release. It does not prove that it will continue to perform that way. Regression testing catches deliberate changes. Recurring validation catches the changes nobody deliberately made.<\/p>\n<\/div><\/div><\/div><\/div>\n<div class=\"et_pb_icon_2 et_pb_icon et_pb_module et_flex_module\"><span class=\"et_pb_icon_wrap\"><span class=\"et-pb-icon\">3<\/span><\/span><\/div>\n<\/div>\n<div class=\"et_pb_column_7 et_pb_column et-last-child et_flex_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide hovergroup\" id=\"trigger-tester\">\n<div class=\"et_pb_blurb_3 et_pb_blurb et_pb_bg_layout_light et_pb_blurb_position_top et_pb_module et_flex_module\"><div class=\"et_pb_blurb_content et_flex_module\"><div class=\"et_pb_blurb_container\"><h3 class=\"et_pb_module_header\">Test the tester<\/h3><div class=\"et_pb_blurb_description\"><p>Automation makes continuous testing possible. But the AI doing the testing has to be tested too. Human review provides an independent check on the evaluator itself.<\/p>\n<\/div><\/div><\/div><\/div>\n<div class=\"et_pb_icon_3 et_pb_icon et_pb_module et_flex_module\"><span class=\"et_pb_icon_wrap\"><span class=\"et-pb-icon\">3<\/span><\/span><\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_section_2 et_pb_section et_section_regular et_flex_section ns-panel\" id=\"define\" style=\"--ns-rail-w:264px\">\n<div class=\"et_pb_row_4 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_8 et_pb_column et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_3 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><p>AI rarely has a simple right-or-wrong output.<\/p>\n<p>Testing therefore starts by translating the intended use and its risks into measurable acceptance criteria: required performance, acceptable error rates, failure behaviour and, where appropriate, probabilistic thresholds.<\/p>\n<p>The depth of testing should follow the consequences of failure. An internal knowledge assistant and a system supporting consequential decisions should not face the same test regime.<\/p>\n<p>This risk-based approach is central to ISO\/IEC TS 42119-2. ISO\/IEC 25059 provides a quality model for AI systems, while ISO\/IEC 23894 addresses AI-specific risk management.<\/p>\n<p><strong>If success cannot be defined, it cannot be tested.<\/strong><\/p>\n<\/div><\/div>\n<\/div>\n<div class=\"et_pb_column_9 et_pb_column et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_4 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><h3>Without a definition of success<\/h3>\n<p>Testing falls back on impressions. Results get argued rather than measured, and \u201cgood enough\u201d is settled after the fact.<\/p>\n<p>Test depth also stops tracking risk: a low-stakes assistant and a system behind consequential decisions end up with the same shallow check.<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_section_3 et_pb_section et_section_regular et_flex_section ns-panel\" id=\"evidence\" style=\"--ns-rail-w:264px\">\n<div class=\"et_pb_row_5 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_10 et_pb_column et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_5 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><p>Development teams need tests they can see and use continuously.<\/p>\n<p>But once a system has repeatedly been optimised against the same cases, those cases become weaker evidence of its actual performance.<\/p>\n<p>A robust test architecture therefore separates:<\/p>\n<ul>\n<li><strong>Development tests<\/strong> used to improve the solution.<\/li>\n<li><strong>Independent validation<\/strong> kept outside the optimisation loop.<\/li>\n<li><strong>Regression tests<\/strong> covering capabilities that must not deteriorate.<\/li>\n<li><strong>Stress and risk tests<\/strong> covering ambiguity, missing information, unusual inputs, adversarial behaviour and application-specific failure modes.<\/li>\n<\/ul>\n<p>Because AI output is variable, important cases also need repetition and variation.<\/p>\n<p>The question is not: can the system give the right answer?<\/p>\n<p>It is: <strong>does it do so reliably enough for the intended use?<\/strong><\/p>\n<\/div><\/div>\n<\/div>\n<div class=\"et_pb_column_11 et_pb_column et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_6 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><h3>What a demo cannot show<\/h3>\n<p>Repeated optimisation against the same cases turns them into part of the build. They stop being evidence of anything.<\/p>\n<p>Without validation kept outside that loop, a convincing demonstration and a reliable system look identical from the outside.<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_section_4 et_pb_section et_section_regular et_flex_section ns-panel\" id=\"baseline\" style=\"--ns-rail-w:264px\">\n<div class=\"et_pb_row_6 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_12 et_pb_column et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_7 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><p>AI systems can change without a conventional software release.<\/p>\n<p>Models are updated. Knowledge bases change. Retrieval produces different context. Prompts and tools evolve. External services change. Users find new ways of interacting with the system.<\/p>\n<p>Known changes should trigger regression testing.<\/p>\n<p>But testing also needs to run independently of releases to establish that production performance remains within the accepted range.<\/p>\n<p>The NIST AI Risk Management Framework calls for testing before deployment and regularly during operation. ISO\/IEC TS 42119-2 likewise treats AI testing as a lifecycle activity and recognises continuous testing as a response to changing system behaviour.<\/p>\n<p><strong>Production is not where testing stops. It is where continuous validation begins.<\/strong><\/p>\n<p><a href=\"\/ns-lab-the-end-of-done\/\">Read: The end of \u201cdone\u201d<\/a><\/p>\n<\/div><\/div>\n<\/div>\n<div class=\"et_pb_column_13 et_pb_column et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_8 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><h3>Why a baseline decays<\/h3>\n<p>Go-live records how the system behaved on one day, with one version of every model, prompt, index and service behind it.<\/p>\n<p>All of those keep moving, and most move without a release. Recurring validation is what keeps the baseline describing the system that is actually in production.<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_section_5 et_pb_section et_section_regular et_flex_section ns-panel\" id=\"tester\" style=\"--ns-rail-w:264px\">\n<div class=\"et_pb_row_7 et_pb_row et_flex_row\">\n<div class=\"et_pb_column_14 et_pb_column et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_9 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><p>Human in the loop does not mean asking people to inspect every AI output.<\/p>\n<p>Humans define evaluation criteria, establish reference cases, investigate material failures and review situations where context or expert judgement matters.<\/p>\n<p>Automation provides the scale required to execute those tests continuously.<\/p>\n<p>But the evaluator also needs oversight.<\/p>\n<p>Cases identified as uncertain or problematic should be escalated for review. Just as importantly, random samples of outputs judged to be successful should also be reviewed by humans.<\/p>\n<p>That distinction matters. Reviewing flagged cases tells us whether the evaluator\u2019s alerts are useful. Random sampling can reveal what it failed to detect.<\/p>\n<p>Human review therefore does more than supervise the AI being tested. <strong>It continuously validates the testing system itself.<\/strong><\/p>\n<\/div><\/div>\n<\/div>\n<div class=\"et_pb_column_15 et_pb_column et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone et_flex_column_24_24_tabletWide\">\n<div class=\"et_pb_text_10 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><h3>Why random samples<\/h3>\n<p>Reviewing flagged cases shows whether the evaluator\u2019s alerts are useful. It says nothing about what the evaluator never flagged.<\/p>\n<p>Sampling the outputs it judged successful is the only way to see that blind spot \u2014 so the tester is measured as deliberately as the system under test.<\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_section_6 et_pb_section et_section_regular et_block_section\" id=\"automated-testing\">\n<div class=\"et_pb_row_8 et_pb_row et_flex_row ns-prose\">\n<div class=\"et_pb_column_16 et_pb_column et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough et_flex_column_24_24 et_flex_column_24_24_tablet et_flex_column_24_24_phone et_flex_column_24_24_tabletWide ns-emboss\">\n<div class=\"et_pb_text_11 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><h2>Automated testing: a stable judge for a moving target<\/h2>\n<p>Continuous AI testing requires automated evaluation.<\/p>\n<p>But if the evaluator itself changes outside your control, it becomes difficult to tell whether a changed test result reflects the system under test \u2014 or the judge.<\/p>\n<p>The 9senses automated testing suite uses a dedicated, calibrated and version-controlled evaluator that can be deployed locally or within a controlled client environment.<\/p>\n<p>Its model version, weights and evaluation configuration remain fixed until they are deliberately changed and revalidated.<\/p>\n<p>That creates a stable benchmark against which changes in the system under test can actually be measured.<\/p>\n<h3>Controlled evaluation<\/h3>\n<p>The current evaluator is based on Mistral and can be used across GenAI and other deep-learning applications where outputs can be assessed against defined criteria.<\/p>\n<p>The suite can continuously:<\/p>\n<ul>\n<li>execute regression and validation suites;<\/li>\n<li>evaluate outputs against use-case-specific acceptance criteria;<\/li>\n<li>identify deviations and emerging failure patterns;<\/li>\n<li>surface edge cases for expert review;<\/li>\n<li>route uncertain cases into human validation;<\/li>\n<li>send random samples for independent review; and<\/li>\n<li>track performance against a known evaluator version over time.<\/li>\n<\/ul>\n<p>Because the evaluator can run inside the testing environment, sensitive test data does not have to be sent to an external frontier-model provider for evaluation.<\/p>\n<h3>The evaluator is part of the test architecture<\/h3>\n<p>A fixed evaluator does not mean an unquestioned evaluator.<\/p>\n<p>Human-reviewed reference cases, edge-case escalation and random sampling are used to validate how the evaluator itself performs.<\/p>\n<p>When a new model version, new weights or a material configuration change is introduced, the evaluator can be recalibrated and validated before becoming the new testing baseline.<\/p>\n<p><strong>The evaluator does not undergo uncontrolled model drift. When it changes, that change is tested too.<\/strong><\/p>\n<\/div><\/div>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"et_pb_section_7 et_pb_section et_section_regular et_block_section\">\n<div class=\"et_pb_row_9 et_pb_row et_block_row ns-cta\">\n<div class=\"et_pb_column_17 et_pb_column et_pb_column_2_3 et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_text_12 et_pb_text et_pb_bg_layout_light et_pb_module et_block_module\"><div class=\"et_pb_text_inner\"><p>Automation provides the scale. Human review provides the control.<\/p>\n<p><strong>9senses Automated Testing is currently in beta.<\/strong><\/p>\n<\/div><\/div>\n<\/div>\n<div class=\"et_pb_column_18 et_pb_column et_pb_column_1_3 et-last-child et_block_column et_pb_css_mix_blend_mode_passthrough\">\n<div class=\"et_pb_module et_pb_button_module_wrapper et_pb_button_0_wrapper\"><a class=\"et_pb_button_0 et_pb_button et_pb_bg_layout_light et_pb_module et_block_module\" href=\"#contact\" style=\"text-wrap:balance\">Talk to us about testing<\/a><\/div>\n<\/div>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"","protected":false},"author":15,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"_acf_changed":false,"_ns_ir":"","_ns_ir_pending":"","_ns_structural_eyebrow":"","_ns_ir_live":"","_ns_structural_enabled":"","_ns_structural_body_v1":"","n9tr_seo_title_de_DE":"","n9tr_seo_description_de_DE":"","n9tr_seo_title_fr_FR":"","n9tr_seo_description_fr_FR":"","footnotes":""},"class_list":["post-226905","page","type-page","status-publish","hentry"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/pages\/226905","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/users\/15"}],"replies":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/comments?post=226905"}],"version-history":[{"count":7,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/pages\/226905\/revisions"}],"predecessor-version":[{"id":230133,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/pages\/226905\/revisions\/230133"}],"wp:attachment":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/media?parent=226905"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}