> ## Documentation Index
> Fetch the complete documentation index at: https://labs.laer.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Review Run scores

| Confidence Score | Interpretation                                                                                            |
| ---------------- | --------------------------------------------------------------------------------------------------------- |
| 90-100           | High confidence. The model found strong evidence supporting the prediction.                               |
| 60-89            | Moderate confidence. The prediction may be valid, but the model identified some ambiguity or uncertainty. |
| 0-59             | Low confidence. The model found limited evidence or could not reach a definitive conclusion.              |

* Prioritize predictions that may require manual validation.
* Focus on low-confidence results first.
* Identify ambiguous documents.
* Assess the certainty of individual field values or tag classifications.
* Sort and filter review results based on model confidence.

Confidence scores represent the model's level of certainty, not the likelihood that a prediction is correct. A confidence score should be interpreted as a ranking and prioritization aid that helps reviewers determine where additional attention may be needed. Confidence scores are most effective when considered together with the model's reasoning and the review strategy for the matter.

Review Runs provide relevance and confidence scores that help you evaluate AI-generated results and identify predictions that may require additional review. These scores are intended to support reviewer decision-making and prioritization, not replace reviewer judgment.

| Score type                | Text Classification Review Runs | Text Generation Review Runs |
| ------------------------- | ------------------------------- | --------------------------- |
| Relevance score (0.0-1.0) | Yes                             | No                          |
| Confidence score (0-100)  | Yes                             | Yes                         |

### Relevance scores

A relevance score indicates how strongly a document relates to the review criteria defined in a **Text Classification Review Run**. Review Run results display relevance scores on a scale from **0.0 to 1.0**. Higher values indicate stronger relevance to the configured review criteria.&#x20;

Epiq AI generates a relevance score together with supporting reasoning. Reviewers can use relevance scores to compare results and prioritize follow-up review.

### Confidence scores

A confidence score indicates the model's level of certainty for a specific prediction. Confidence scores are displayed on a scale from **0 to 100**. Confidence scores are available for both **Text Classification** and **Text Generation** Review Runs.

#### Confidence scores in Text Generation Review Runs

Text Generation models generate results for the output fields defined in the Review Run configuration.&#x20;

For each output field, Epiq AI returns:

* Field name
* Generated value
* Confidence score
* Reasoning

Each field is evaluated independently. As a result, a document can contain both high-confidence and low-confidence results, depending on the available evidence for each field.&#x20;

For example, if a Review Run contains three output fields, Epiq AI returns three separate generation results, each with its own confidence score and reasoning.&#x20;

Confidence scores are also displayed on the Review Run table and at the document level.

### Relevance and confidence scores in Text Classification Review Runs

Text Classification models evaluate the tags defined in the Review Run configuration.

For each tag, Epiq AI returns:

* Tag name
* Verdict
* Confidence score
* Reasoning

Each tag is evaluated independently. As a result, a document can receive different confidence scores for different classifications. For example, if a Review Run evaluates eight tags, Epiq AI returns eight separate classification results. Text Classification Review Runs also provide a document-level relevance score and a document-level confidence score to help reviewers evaluate and prioritize results.

### Interpret confidence scores

The following ranges provide general guidance for interpreting confidence scores.

| Confidence score | Interpretation                                                                                            |
| ---------------- | --------------------------------------------------------------------------------------------------------- |
| 90-100           | High confidence. The model found strong evidence supporting the prediction.                               |
| 60-89            | Moderate confidence. The prediction may be valid, but the model identified some ambiguity or uncertainty. |
| 0-59             | Low confidence. The model found limited evidence or could not reach a definitive conclusion.              |

These ranges are guidelines only. Use confidence scores to compare and prioritize results rather than as a guarantee of accuracy.

### Use confidence scores during review

Confidence scores can help reviewers:

* Prioritize predictions that may require manual validation.
* Focus on low-confidence predictions.
* Identify ambiguous documents.
* Assess the certainty of individual classifications.
* Sort and filter results based on model confidence.

Because confidence scores are generated for individual predictions, reviewers can focus on specific results that require additional attention rather than treating an entire document as high or low confidence.

### Confidence scores and accuracy

Confidence scores represent the model's level of certainty, not whether a prediction is correct. Use confidence scores as a ranking and prioritization aid when reviewing results. Confidence scores are most effective when considered together with:

* The prediction itself
* The model's reasoning
* The review objectives for the matter

<Info>
  **Note:** Confidence scores support human review but do not replace reviewer judgment. Reviewers should validate predictions according to the requirements and review strategy of their matter.
</Info>
