Skip to main content
The objective of AIDA sandbox environment is to offer an opportunity to explore AIDA’s different features and capabilities on test data. This document summarizes some of AIDA’s capabilities, listed below, and provides a walkthrough through each capability with recommendations on how to explore and evaluate it. AIDA’s key capabilities include:
  1. Conversational Inquiry. The ability to ask factual and review management questions through the chat interface or by email and receive substantive responses and documents.
  2. Conversational Review. Ability to upload written classification instructions (“Model Instructions”), such as a review protocol, subpoena, or document request, and quickly build predictive AI models to identify and organize relevant content.
  3. Knowledge Layer. Ability to explore the Knowledge Layer that AIDA builds upon first ingesting the data, containing entities and facts that AIDA extracts by reading all matter data and utilizing these representations to enhance AIDA capabilities across the platform including in Conversational Inquiry and Conversational Review.
The plan for AIDA sandbox evaluation is as follows:
  1. Conversational Inquiry
    1. Asking questions
      1. Factual Questions
      2. Chronology Questions
      3. Review Management Questions
    2. Inspecting evidence
  2. Conversational Review
    1. Creation of Model Instructions
    2. Building predictive models
    3. Exploring built model predictions
    4. Annotating documents
    5. Retraining models
  3. Knowledge Layer
    1. Inspecting Knowledge Layer
    2. Editing Knowledge Layer
There are two test cases that are available as part of the sandbox:
CaseSizeReview DataFactual Conv. InquiryReview Conv. Inquiry
Enron640KYesYesYes
Mallinckrodt1.4MNoYesNo
The two cases differ in their content and have specific nuance to be aware of during evaluation: - Enron: Contains original data from an NIST-sponsored Text Retrieval Conference Study. Appropriate for both factual and review management qualitative and quantitative evaluation as it contains both Model Instructions (as a review protocol), and reviewer tags created as part of the industry TREC study. - Mallinckrodt: Contains only the original data from the UCSF collection. This original data set does not contain a review protocol or reviewer tags, and also suffers from poor quality OCR which will impact data quality especially as extracted and stored as values in the Knowledge Layer. This dataset would be appropriate for qualitative evaluation given these considerations.