Priority Sampling
The goal is to prioritize documents that are most likely to be relevant based on model predictions. The sampling process is: The process:- 90% of the batch comes from the highest-scoring documents: sort the documents in descending order, apply clustering on the top-ranking documents, and select the centroids to form the sample batch. This is to ensure some degree of diversity in the documents.
- 10% of the batch comes from a pure random sample.
Coverage Sampling
The goal is to ensure diverse areas of the document space are covered, avoiding bias toward highly scored documents. The process:- 30% of the batch comes from documents with scores above 0.5. Documents are selected as centroids of document clusters to ensure diversity.
- 30% of the batch comes from documents with scores below 0.5. Documents are selected as centroids of document clusters to ensure diversity.
- 30% of the batch comes from documents with scores around 0.5. Documents are selected as centroids of document clusters to ensure diversity.
- 10% of the batch comes from a pure random sample.
Balanced Sampling
The goal is to balance the review process by combining priority and coverage strategies. The process:- 70% of the batch is selected from documents with scores above 0.7. Documents are selected as centroids of document clusters to ensure diversity.
- 20% of the batch comes from documents with scores below 0.7. Documents are selected as centroids of document clusters to ensure diversity.
- 10% of the batch comes from a pure random sample.