How is your GenAI application doing?
9senses Generative AI Audit
We have all come across Generative AI applications designed to help us – from customer-facing chatbots and internal knowledge assistants to retrieval systems, summarization tools and Copilot implementations. They seem to be everywhere by now.
Sometimes they work surprisingly well by delivering useful answers or results within a few seconds and saving us from spending half an hour in a call center queue, navigating tons of subpages, or manually searching through information.
On the other hand, we are often stuck with an AI application that does not understand what we want, returns irrelevant or unreliable results, and leaves us feeling that we have wasted valuable time on a solution that could not care less about our problems.
This is why 9senses created its GenAI Audit. Our Level 1 GenAI Audit is a standardized assessment tool to measure the performance of your GenAI application and compare it with best-in-class applications. Using a carefully developed and tested methodology within the 9senses GenAI Audit Framework, we review your application, assess its performance, and identify where it can be improved.
Level 1 Audit
An independent, external review of your GenAI application's observable performance. No internal access to your technical environment required. Delivered within 5 business days.
Level 2 Audit
A tailored deep-dive analysis, for example of technical architecture, retrieval systems, governance, compliance, and business value – recommended when critical issues are identified in Level 1.
FAQs - Frequently Asked Questions
What is a GenAI Audit?
A 9senses GenAI Audit is an independent assessment of how well a Generative AI application performs in real-world use. It evaluates the application from the user's perspective and whether it reliably achieves its intended purpose.
The Level 1 GenAI Audit is a standardized Black-Box Audit based on the 9senses GenAI Audit Framework. We test the application using agreed use cases and assess Response Quality, Speed, User Interface, and Dialog Quality. The result is a structured report showing how the application performs, where its strengths and weaknesses lie, and where improvements should be prioritized.
What types of GenAI applications can you audit?
We can audit a broad range of Generative AI applications, including customer-facing chatbots, internal AI assistants, knowledge and retrieval systems, RAG implementations, summarization tools, Microsoft Copilot implementations, and other LLM-based applications.
The underlying principle is the same: we define realistic use cases based on how the application is intended to be used and derive concrete test cases to assess the quality of the results it produces.
The audit can therefore be applied to applications serving customers, employees, or specialist user groups, whether the application is conversational or embedded in a broader workflow.
What does the Level 1 GenAI Audit include?
The Level 1 GenAI Audit is a Black-Box Audit of observable performance. We interact with the application as a real user would, without requiring access to its underlying model, prompts, retrieval architecture, source systems, or other internal technical components.
The audit begins with a briefing and an agreed set of use cases. We derive concrete test cases, execute the agreed test set, and assess the application across the 9senses GenAI Audit Framework: Response Quality, Speed, User Interface, and Dialog Quality. A Translation Test can be added when relevant.
You receive a structured report with scores, findings, and prioritized recommendations for improvement.
How does the GenAI Audit detect hallucinations and other risks?
Hallucinations – the creation of invalid or invented content – pose a serious reputational and compliance risk. The audit includes targeted hallucination stress testing as part of Response Quality.
We deliberately introduce invalid references, misspellings, and ambiguous prompts to assess whether the application fabricates information or requests clarification. The evaluation examines entity validation, grounding behavior, and escalation logic.
Where Variance Testing is selected, we also repeat and rephrase prompts to assess how consistently the application responds when users express essentially the same request in different ways.
Does the GenAI Audit assess compliance (EU AI Act, GDPR)?
Level 1 includes a preliminary external review of observable compliance and transparency indicators, including AI disclosure, GDPR-relevant interface elements, and accessibility.
A full regulatory and governance analysis – including documentation and architecture review – can be part of a tailored Level 2 Audit.
What information do you need, and how long does the audit take?
We need enough context to understand what the application is intended to do, who uses it, and what good performance looks like.
For the standard audit, you provide a short briefing and 3–5 relevant use cases representing realistic situations the application should be able to handle. If you do not book Use Case Creation, we provide a form for entering your use cases. If you book Use Case Creation, 9senses develops 3–5 realistic, business-relevant use cases based on your briefing and sends them to you for review before testing.
For publicly accessible applications, we need access to the live application interface. For login-protected applications, we also need the required test access. Where the audit depends on internal knowledge data or other defined source material, we need access to the application interface and to the knowledge base or source material it uses so that we can validate the relevance and grounding of its responses.
The Level 1 GenAI Audit is completed within 5 business days after we have received the briefing, agreed use cases, and any required access information. If you book Use Case Creation, allow 2 additional business days for delivery and review of the proposed use cases.
How do you ensure confidentiality?
All audit work and deliverables are handled under confidentiality. Reports and findings are shared only with the customer. Numeric audit results are used for best-in-class benchmarking.
What options can I add to the Level 1 GenAI Audit?
The Level 1 GenAI Audit can be tailored to your application’s setup and business context. In addition to the core audit, you can select the following options:
-
Use Case Creation
9senses defines 3–5 realistic, business-relevant use cases based on your briefing and sends them to you for review before testing. Select this option if you do not yet have structured evaluation cases. -
Open Internet Search
For applications that retrieve information from external websites or search engines. This option evaluates source consistency, grounding behavior, and increased hallucination exposure. -
Login-Protected Access
For applications accessible only behind login, such as customer portals or employee systems. This enables evaluation within authenticated environments and role-specific flows. -
Variance Testing
Tests the same initial user prompt twice to identify variance and includes a rephrased version of the initial prompt to stress-test response consistency. -
Translation Test
Includes structured verification of one additional language, focusing on consistency, language-switching behavior, and translation accuracy. -
Executive Briefing
A structured management-level walkthrough of findings. We translate audit results into decision priorities, risk implications, and concrete next steps.
These options allow you to adapt the GenAI Audit to your application’s setup, risk exposure, and governance requirements.
What is the difference between Level 1 and Level 2, and when should I consider Level 2?
A Level 1 GenAI Audit is a Black-Box Audit focused on observable, user-facing performance. It does not require access to the underlying technical environment and provides a fast, independent view of how well the application performs.
A Level 2 Audit is an Open-Book Audit tailored to the issues or questions that require deeper investigation. It can examine the application's technical architecture, retrieval/RAG systems, governance, compliance, risk management, and business value modeling based on detailed internal insights.
Level 2 is recommended when a Level 1 Audit identifies issues that require deeper investigation - for example, a score below 3.5 in Response Quality or Dialog Quality for an application with high user visibility. It is also appropriate for strategically important applications or systems serving defined user groups, such as customers or employees, where business or regulatory risk is material.
