Overview
The EAIDA Privilege Workflow is a comprehensive system designed to streamline the privilege review process for legal document analysis. This workflow automates the extraction, normalization, classification, and privilege log generation for document collections, significantly reducing manual effort while maintaining high accuracy.Workflow Architecture
The EAIDA Privilege workflow consists of six interconnected components that build upon each other:Workflow Dependencies
These components have dependencies where the quality of output from previous steps impacts the current step. Understanding these dependencies is crucial for effective quality control and troubleshooting.Getting Started
Prerequisites
Before beginning your privilege workflow, ensure you have:- Dataset Creation: A dataset containing documents ready for privilege analysis
- Libraries Setup: Upload known attorney and domain lists to Libraries for consistent entity treatment
- Collection Access: Appropriate permissions to the target collection
Creating a New Workflow
Access Workflows
- Sign in to EAIDA
- Navigate to your workspace (e.g., “AI lab demo”)
- Click Collections
- Select your target collection (e.g., “Edison set”)
- Click Protocols
- Navigate to Workflows
Select Workflow Template
- Choose from pre-built templates or create a custom workflow
- Templates indicate the number of components included
- Workflow configurations can be saved as templates for future use
Workflow Navigation
The privilege workflow interface contains six main screens:Step-by-Step Workflow Guide
Importing Data
Purpose: Populate your workflow with the document dataset to be analyzed.Import Data (Start)
- Click Input Data from the workflow screen
- Click Add Data
- Select Import an existing dataset
- Click Select Dataset
- Choose your dataset (e.g., “Sample set”)
- Click Import
Key Points
- You need a Dataset in the Collection you’re targeting
- To create a Dataset, drill into the target Collection and create it in the Collection screens
- Once imported, documents will appear under “Documents”
- You can add multiple datasets to the workflow with additional imports
- View documents for each import by clicking the Import card, or view the entire workflow population
Configuration
Access the Configuration screen to set up each component before deployment. Each component below outlines purpose, what it does, configuration steps, and notes.Metadata Extraction
Purpose: Extract targeted metadata from source documents, including participants from metadata and nested email headers.Configure Metadata Extraction
- Click Configuration from the main menu
- Select Metadata Extraction
- Review and confirm default fields
- Add new fields or composite fields as needed
What This Component Does
- Extracts participants from document metadata
- Parses nested email headers from message bodies
- Populates designated fields for the Name Normalization component
- Confirm default extraction fields
- Define custom fields for specific metadata
- Create composite fields for complex data structures
Name Normalization
Purpose: Standardize participant and organization names, consolidating name variations into Actor (Person) and Domain (Organization) entities.Configure Name Normalization
- From Configuration, click Name Normalization
- Confirm default output fields for Actors and Domains
- Review normalized format settings
- Create new Composite fields if you need formats beyond the standard
What This Component Does & Key Concepts
- Takes participants and organizations from Metadata Extraction
- Normalizes their names to standard formats
- Consolidates name variations (e.g., “John Smith”, “J. Smith”, “Smith, John”)
- Creates Actor entities (people) and Domain entities (organizations)
- Actors: Term for Person entities in the Privilege Workflow
- Domains: Term for Organization entities in the Privilege Workflow
These are distinct from entities identified automatically in the Knowledge Layer
- Set output field formats for Actors and Domains
- Define composite fields for custom name formats
- Confirm input fields from Metadata Extraction step
Entity Labeling
Purpose: Assign legal roles to identified Actors and Domains based on their function in the matter.Configure Entity Labeling
- From Configuration, select Entity Labeling
- Review and edit instructions for Persons (Actors)
- Review and edit instructions for Organizations (Domains)
- Configure General Instructions
- Set up Tag Dependencies
Default Person (Actor) Tags
Default Organization (Domain) Tags
Important Configuration Notes & Tag Dependencies
- Fill in client and counterparty names in the instructions for accurate classification
- Customize tags and instructions according to the specific matter requirements
- Roll up granular categories into general categories
- Create mutually exclusive category rules
- Entities tagged as “Client Law Firm” OR “Opposing Law Firm” should also be identified as “Lawyers”
- Entities classified as “Law Firm” cannot also be “Third Party”
Metadata Normalization
Purpose: Apply normalized metadata and entity records to documents in the workflow.Configure Metadata Normalization
- From Configuration, click Metadata Normalization
- Review settings (this step has no configuration options)
What This Component Does
- Associates each participant’s cleaned-up name with relevant documents
- Links entity records to documents where they appear
- Applies all normalization from previous steps
Privilege Classification
Purpose: Classify each document by privilege type using AI models and entity role information.Configure Privilege Classification
- From Configuration, select Privilege Classification
- Click to edit classification instructions
- Review and customize Tag definitions
- Configure General Instructions
- Set up Tag Dependencies
- Configure Model settings (Positive/Negative Labels, Train/Exclude datasets)
- Configure Metrics and Error Categories
Default Classification Tags
Metadata Field Population, Dependencies & Model Config
- Entity Tags (Metadata)
- Entity Tags (Extracted Metadata)
- Entity Tags (Mentioned)
- Parent Entity Tags (Metadata)
- Make tags mutually exclusive (e.g., a document cannot be both “Not Privileged” and “Privileged”)
- Create hierarchical classification structures
- Set Positive and Negative Labels
- Define Train and Exclude Train datasets
- Configure threshold settings
- Customize tags and instructions based on the specific matter
- Set up clear tag dependencies to avoid conflicting classifications
- Use Error Categories in Metrics configuration for quality control
Privilege Log Generation
Purpose: Generate privilege log descriptions using text generation models. Prerequisite: Privilege Classification must be completed before this step runs.Configure Privilege Log Generation
- From Configuration, select Privilege Log Generation
- Click to edit privilege log instructions
- Review and customize field definitions
- Configure Metrics and Error Categories
Default Privilege Log Fields
Customization Tips
- Adjust field instructions to match your organization’s privilege log format
- Ensure descriptions are sufficiently detailed but don’t reveal privileged information
- Customize examples to reflect typical document types in your matters
Deploying the Workflow
Once all components are configured, deploy the workflow to begin processing. Deployment Options: Option 1: Deploy All at Once- From the Configuration screen, click Deploy Workflow
- Click Deploy to confirm
- All configured components will run in sequence
- Navigate to the Flow screen
- Run each component individually in order
- Review output before proceeding to the next step
Monitoring and Reviewing Results
Using the Flow Screen
The Flow screen is your central hub for monitoring progress and reviewing results. Flow Screen Capabilities:- Run workflow components
- Monitor what is currently running
- View output results
- Export output for correction
- Import corrected output (append or overwrite)
- View or download output from previous runs
Metadata Extraction Results
Viewing Results:- Click Flow from the main menu
- Select Metadata Extraction
- Choose viewing mode:
- Document Level: View metadata for individual documents
- Thread Level: View metadata organized by email threads
- Verify all documents have extracted metadata
- Check that nested email headers are parsed correctly
- Confirm all participants are identified
- Count of documents where metadata was extracted
- Count of entities identified per document
- Documents missing expected metadata
Name Normalization Results
Viewing Results:- From Flow, click Name Normalization
- View output for Actors and Domains separately
- Results are normalized but NOT consolidated
- Each email address appears as a separate entry
- One individual may have multiple entries (one per email variant)
- Add and correct entries individually through the interface
- Export full results for bulk corrections
- Import corrected entries to update the workflow
- Missing first names or last names
- Missing organization information
- Non-standard characters in normalized output
- Inconsistent handling of honorifics (e.g., “Dr”)
- Improper parsing of combined addresses (e.g., “Joe and Jane Smith”)
Entity Labeling Results
Viewing Results:- From Flow, click Entity Labeling
- View output for Person and Organization separately
- Results show CONSOLIDATED entities
- One entry per unique individual/organization
- Entity tags assigned based on configured instructions
- Add and correct entries individually
- Export full results for bulk corrections
- Import corrected entries
- Access the Metrics screen to view statistics about identified entities
- Review entity tag distributions
- Identify potential classification issues
- Verify attorney and counsel identifications
- Check client and opposing party classifications
- Review third-party designations
Metadata Normalization Results
Viewing Results:- From Flow, click Metadata Normalization
- Review formatted names for all extracted metadata
- Confirm all documents have normalized metadata entries
- Verify entity records are correctly linked to documents
Privilege Classification Results
Viewing Results:- From Flow, click Privilege Classification
- Choose viewing mode:
- Document Level: Individual document classifications
- Thread Level: Thread-based classifications
- Create Review Batches: Organize documents for manual review
- Review Assignments: Assign documents to reviewers
- Error Analysis: identify potential misclassifications and edge cases
- Model Metrics: review classification performance charts, adjust cutoff thresholds, view precision/recall metrics
- Review documents at classification boundaries
- Verify “Potentially Privileged” designations
- Check for conflicting tags
- Validate “Not Privileged” classifications where attorneys are present
Privilege Log Generation Results
Viewing Results:- From Flow, click Privilege Log Generation
- Review generated privilege log fields
- Export results for final review and publication
- Create review batches for privilege log QC
- Assign privilege log entries for review
- Access Metrics screen for quality statistics
- Export complete privilege log
- Export by privilege category
- Export for specific date ranges
- Verify all privileged documents have descriptions
- Check for missing field values
- Review description fluency and accuracy
- Ensure no privileged information is disclosed in descriptions
- Validate document categories match file types
- Check for sentence fragments or incomplete entries
- Blank Privilege Communication Type fields
- Missing Privilege Descriptions for privileged documents
- Inconsistent formatting
- Documents with no tag values
- Contradictory tags
Correcting and Refining Results
For Name Normalization, Entity Labeling, and other components, follow the Export → Correct → Import process.Export Results
- Navigate to the component in Flow
- Click the export option
- Download the results file
Make Corrections
- Edit the exported file using appropriate tools
- Correct entity classifications
- Fix normalization errors
- Add missing information
Import Corrections
- Return to the component in Flow
- Click the import option
- Select your corrected file
- Choose import mode:
- Append: Add new entries without replacing existing ones
- Overwrite: Replace existing entries with corrected versions
Re-run Downstream Components
- After importing corrections, consider re-running downstream components
- This ensures corrections propagate through the workflow
Output Screen
The Output screen provides a consolidated view of workflow results across all components. Accessing Output:- Click Output from the main menu
- View aggregated results from all workflow stages
- Final review before export
- Cross-component validation
- Summary statistics
Runs History
The Runs screen maintains a complete history of workflow executions. Accessing Runs:- Click Runs from the main menu
- View chronological list of workflow runs
- Timestamp of each run
- Components executed
- Run status (completed, failed, in progress)
- User who initiated the run
- Troubleshooting workflow issues
- Tracking workflow progress over time
- Comparing results between runs
- Audit trail for quality control
Best Practices
Pre-Workflow Preparation
- Clean Your Dataset
- Remove duplicates before importing
- Ensure documents are properly loaded
- Verify metadata is present
- Prepare Your Libraries
- Upload known attorney lists
- Include domain lists for key organizations
- Update lists as new entities are discovered
- Customize Instructions
- Fill in client and counterparty names in Entity Labeling
- Adjust classification instructions for matter-specific requirements
- Customize privilege log descriptions for your organization’s standards
Configuration Best Practices
- Start with Default Settings
- Use default configurations as a baseline
- Make incremental adjustments
- Document any changes for future reference
- Test with a Sample
- Run the workflow on a small sample first
- Validate results at each stage
- Adjust configurations before processing the full dataset
- Set Up Tag Dependencies Carefully
- Ensure mutually exclusive tags are properly configured
- Create logical hierarchies for roll-up tags
- Test dependency rules with sample data
Quality Control
- Review at Each Stage
- Don’t wait until the end to check results
- Validate output at each component before proceeding
- Correct errors early to prevent propagation
- Use Metrics Effectively
- Monitor entity counts and distributions
- Review classification metrics regularly
- Set up error categories for systematic QC
- Manual Review Priorities
- Focus on “Potentially Privileged” documents
- Review edge cases at classification boundaries
- Validate entries with missing information
Iterative Improvement
- Learn from Each Run
- Document issues encountered
- Update instructions based on errors
- Refine tag definitions over time
- Update Libraries
- Add newly identified attorneys to libraries
- Update domain lists with new organizations
- Maintain libraries across matters
- Save Successful Configurations
- Create templates from well-configured workflows
- Document custom settings
- Share best practices across your team
Evaluation Metrics by Component
Metadata Extraction
Key Metrics:- Count of documents where metadata was extracted
- Count of entities per document
- Number of documents with missing metadata
- Nested headers not parsed correctly
- Missing participants from headers
- Incomplete extraction from complex documents
Name Normalization
Key Metrics:- Count of identified entities
- Count of incorrect entities
- Number of missing first names
- Number of missing last names
- Number of missing organizations
- Email strings with different markers not consolidated
- Combined addresses not parsed correctly
- Inconsistent handling of honorifics
- Non-standard characters in output
Entity Labeling
Key Metrics:- Count of Attorney + Client Counsel
- Count of Attorney + Legal Staff
- Count of Attorney + Opposing Counsel
- Count of each label standalone
- Count of labeling errors
- Verify attorney identifications are complete
- Check for misclassified entities
- Ensure client and opposing counsel are correctly identified
Metadata Normalization
Key Metrics:- Count of documents missing normalized entries
- Count of linkage errors
Privilege Classification
Key Metrics:- Precision and recall compared to human labels
- Distribution of classifications across categories
- Count of documents at classification boundaries
- Compare with prior review data if available
- Sample-based manual review
- Focus on “Potentially Privileged” category
Privilege Log Generation
Key Metrics:- Count of entries missing information
- Count of disfluent entries
- Count of entries needing manual correction
- Count of documents with no tag value
- Verify Privilege Communication Type is filled
- Check Doc Category matches File Type
- Review fluency of Privilege Descriptions
- Ensure no privileged content is disclosed in descriptions
Troubleshooting
Common Issues and Solutions
Issue: Metadata Extraction Missing Participants
Possible Causes:- Complex nested email headers
- Non-standard metadata format
- Corrupted document files
- Review default field configurations
- Add custom fields for non-standard metadata
- Check source document quality
Issue: Name Normalization Not Consolidating Variants
Possible Causes:- Different email formatting (brackets vs. angle brackets)
- Special characters in names
- Inconsistent name formats in source
- Export and manually correct entity mappings
- Update normalization rules
- Add known variants to Libraries
Issue: Entity Labeling Misclassifying Roles
Possible Causes:- Insufficient or ambiguous instructions
- Missing client/counterparty names in instructions
- Conflicting tag definitions
- Customize tag instructions with specific matter details
- Fill in client and counterparty names
- Set up clear tag dependencies
- Export, correct, and re-import entity classifications
Issue: Privilege Classification Inconsistent Results
Possible Causes:- Poor quality entity labeling from previous step
- Insufficient training data
- Overly broad or narrow classification instructions
- Review and correct entity labeling results
- Adjust classification threshold
- Refine tag instructions
- Provide additional training examples
- Use tag dependencies to resolve conflicts
Issue: Privilege Log Missing Descriptions
Possible Causes:- Documents not properly classified as privileged
- Incomplete privilege classification
- Missing entity information
- Review privilege classification results
- Ensure all privileged documents are properly tagged
- Correct entity labeling upstream
- Customize Privilege Description instructions
Appendix: Using Libraries for Enhanced Accuracy
What are Libraries?
Libraries in EAIDA are curated lists of known entities (attorneys and domains) that give the workflow components a head start on entity identification and classification.Setting Up Libraries
Before Starting Your Workflow:- Compile lists of known attorneys for the matter
- Create lists of relevant domains (law firms, client organizations, opposing parties)
- Upload these lists to Libraries in EAIDA
Benefits of Using Libraries
- Improved Accuracy: Known entities are consistently identified and classified
- Reduced Manual Correction: Fewer errors to fix downstream
- Faster Processing: Pre-identified entities speed up normalization and labeling
- Consistency Across Matters: Maintain entity information across multiple workflows
Maintaining Libraries
- Update libraries as new entities are discovered during review
- Add entities from corrected workflow output back to libraries
- Share libraries across matters involving the same parties
- Periodically review and clean up library entries
Glossary
- Actor: The Privilege Workflow term for a Person entity; represents an individual participant in documents.
- Domain: The Privilege Workflow term for an Organization entity; represents a company, law firm, or other organizational participant.
- Entity Labeling: The process of assigning legal roles (e.g., Attorney, Client Counsel) to identified participants.
- Metadata Extraction: The automated process of extracting participant and organizational information from document metadata and headers.
- Name Normalization: The process of standardizing name variations and consolidating them into unique entities.
- Libraries: Curated lists of known attorneys and domains used to improve entity identification accuracy.
- Tag Dependencies: Rules that define relationships between classification tags, such as hierarchies or mutual exclusions.
- Composite Fields: Custom field definitions that combine multiple metadata elements.
Additional Resources
Need Help?
If you encounter issues not covered in this documentation:- Review the Runs history to identify where the workflow failed
- Check component-specific metrics for detailed diagnostics
- Export results for manual review and correction
- Contact your EAIDA administrator or support team
Future Enhancements
Planned features for future releases:- Updating and correcting output with import overlay of corrected annotations
- Enhanced QC workflows
- Advanced metrics for Entity Labeling, Privilege Classification, and Privilege Log Generation
- Standalone privilege log component execution
Conclusion
The EAIDA Privilege Workflow provides a powerful, automated approach to privilege review and log generation. By following this guide and implementing best practices, you can:- Significantly reduce manual review time
- Improve consistency and accuracy in privilege determinations
- Generate high-quality privilege logs efficiently
- Maintain comprehensive audit trails