Skip to main content
Metrics are available on the PII Detection component. Select the component from the Flow tab, then select the Metrics view.

Analyzing detection errors

Before the Prediction Overview and Error Analysis tabs have data, you need to run a feedback batch. This instructs the system to compare the AI’s determinations against your reviewers’ decisions and generate explanations for any mistakes. To run feedback, select the documents you want to evaluate and click Analyze Errors. The process runs in the background. Once complete, all three tabs will reflect the results. You can re-run feedback any time to refresh the data when additional review work has been completed.

Model Metrics

This tab shows how the detection model is performing based on a set of documents that have already been reviewed and coded. Summary cards at the top:
  • Total coded documents: the number of documents that have been manually reviewed.
  • Total coded positive: how many of those were confirmed as containing PII or PHI.
  • Total coded negative: how many were confirmed as containing neither.
Relevance Score chart — a bar chart showing how many documents fall into each confidence range. Bars to the left of the cutoff threshold are classified as negative (does not contain PII); bars to the right are classified as positive (contains PII). You can drag the threshold slider to explore how changing the cutoff affects your results. For privacy review the threshold is a scope decision, not just a tuning one. Lowering it pulls more documents into the reportable population and increases the volume a reviewer has to confirm; raising it does the opposite. Set it deliberately for the matter.

Prediction Overview

This tab gives you a high-level summary of what the AI determined across your entire document set, and how much of that work has been reviewed by a human. Summary cards:
  • Total Documents: the total number of documents the AI has scored.
  • Contains PII: documents the AI determined to contain PII or PHI.
  • Does Not Contain PII: documents the AI determined to contain neither.
Tags Analysis chart: a stacked bar chart breaking down determinations by element type (for example, Social Security number, date of birth, medical record number). For each element:
  • The green portion shows documents that have been manually reviewed.
  • The gray portion shows documents that have not yet been reviewed.
This chart helps you understand which elements still have unreviewed determinations, and which elements are driving the population.

Error Analysis

This tab is where you can dig into the AI’s mistakes, understanding why errors occurred.

Filters

Three filters at the top let you focus on a specific subset of results:
  • Tags: narrow down to a specific element or determination.
  • Error Type: filter by:
    • All Errors: show everything
    • False Positives: documents the AI flagged as containing sensitive information, but a reviewer said did not
    • False Negatives: documents the AI missed (flagged as clean, but a reviewer found sensitive information)
  • Confidence Level: focus on determinations the AI was Low (0–60%), Medium (60–80%), or High (80–100%) confident about.
False negatives are the errors that matter most in a privacy review — an individual missed here is an individual not notified. Review them first.

Error Categories Distribution

A pie chart showing which error categories appear most frequently across all mistakes. Each slice represents a category of reason the AI got it wrong. Hover over any slice to see the exact count and percentage. Error categories can be customized in the Configuration tab.

Error Analysis Table

A detailed row-by-row table of each document the AI was evaluated on. Columns include: You can apply additional filters using the filter bar.

Workflow reports

Metrics measure how the model is performing. The Output tab carries the reports that describe what the review found:
  • Prediction output — documents reviewed, documents containing PII, documents containing PHI, affected individuals, and a breakdown of each sensitive data element by the number of individuals and documents it appears in.
  • Exceptions — a reconciliation of documents submitted against documents that produced a result, with the shortfall broken down by cause and a recommended action for each.
Use metrics to decide whether the results are good enough, and the reports to communicate what the results are.