Help improve enterprise AI data pipelines by testing data quality, PII and PHI de-identification, and ML evaluation systems. This remote contract role offers 20+ hours per week for experienced data and ML engineers.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect skilled contributors with work that helps shape modern AI systems, while providing a place to build a lasting professional profile and portfolio.
OpenTrain AI is hiring and contracting for this role. You can apply through OpenTrain and manage your AI training career from one account.
- Remote work available worldwide
- Contractor and part-time engagement
- Work in English
- Expected commitment of 20+ hours per week
About AI Training and Data Quality Work
AI systems depend on reliable data prepared, tested, and reviewed by people. Data quality and privacy specialists help ensure that the information used to train and evaluate models is accurate, consistent, and appropriately protected.
In this role, your work will support AI pipelines by testing enterprise data, evaluating de-identification systems, and identifying privacy risks across raw, processed, and sanitized datasets.
- Contribute to the human quality and compliance layer behind AI systems
- Work with enterprise sources, multilingual entities, OCR-affected content, and unusual document formats
- Use technical evaluation to improve the reliability and safety of AI data workflows
The Role
OpenTrain is seeking a Data Quality and PII Compliance Engineer to improve the reliability, privacy, and compliance of enterprise data pipelines used in AI systems. The role combines data engineering, ML evaluation, test automation, and compliance-focused quality assurance.
You will focus on testing PII and PHI de-identification systems, assessing whether sensitive information has been removed accurately, and measuring the performance of NER and other ML-based systems with labeled datasets.
- Subject matter focus: PII De-identification Quality Evaluation
- Engagement: Contractor, part time
- Workload: 20+ hours per week
- Location: Worldwide and remote
- Experience level listed: Entry level
What You'll Do
You will design, automate, and maintain quality and privacy checks across enterprise data pipelines. The work includes technical analysis, adversarial testing, human audits, compliance documentation, and operational monitoring.
- Design validation suites for schema checks, completeness, drift detection, and reconciliation across pipeline stages
- Analyze coherence, domain coverage, consistency, and depth across enterprise data sources and connectors
- Build adversarial datasets for de-identification testing, including edge cases and obfuscated identifiers
- Evaluate NER and other ML-based systems using precision, recall, F1 score, leak rates, and false-negative analysis
- Assess non-deterministic ML systems with labeled datasets and appropriate evaluation metrics
- Investigate privacy leakage across raw, processed, and sanitized data
- Implement CI/CD regression gates for data quality and privacy checks
- Conduct sampling-based human audits and maintain compliance audit trails
- Develop monitoring and reporting for pipeline quality, de-identification performance, and compliance risk
Required Skills and Experience
This role requires strong technical experience in data engineering, ML engineering, data quality, or a related field. The listing is marked entry level, but the stated requirements include 5+ years of relevant experience.
- 5+ years of experience in data engineering, ML engineering, data quality, or a related field
- Strong Python and SQL skills for validation, test automation, and ML evaluation
- Knowledge of PII, PHI, and de-identification concepts
- Familiarity with HIPAA Safe Harbor, GDPR, or LGPD
- Experience evaluating NER or other ML-based systems with labeled datasets
- Understanding of precision, recall, F1, leak-rate, and false-negative metrics
- Experience evaluating non-deterministic ML systems
- Experience integrating data-quality and privacy checks into CI/CD pipelines
- Familiarity with pytest, Great Expectations, Pandera, or similar tools
- Familiarity with GCP services such as BigQuery, Google Cloud Storage, or Cloud Run Jobs
- Strong analytical, debugging, and communication skills
Who Should Apply
This opportunity is suited to professionals who can combine engineering discipline with careful privacy and model evaluation. It may be a strong fit if you have worked with enterprise data pipelines, labeled evaluation datasets, automated testing, or de-identification quality assurance.
- Data engineers focused on pipeline reliability and automated validation
- ML engineers experienced in evaluating NER or related systems
- Data quality specialists who understand schema, drift, completeness, and reconciliation checks
- Privacy and compliance-focused QA professionals with technical testing experience
- Engineers comfortable analyzing edge cases, debugging failures, and documenting audit evidence
How OpenTrain Supports Your Career
AI training and data-labeling work is a fast-growing part of the technology industry. Contributors help build better AI by preparing examples, reviewing model behavior, evaluating outputs, and applying specialized expertise to complex datasets.
OpenTrain gives you one place to discover relevant opportunities, build a credible profile, and develop a portfolio of AI training experience. Creating an OpenTrain account is free, and you can apply to this role in minutes.
- Build a durable portfolio of AI training and evaluation work
- Showcase technical experience in data quality, privacy, and ML evaluation
- Find flexible remote opportunities that match your skills
- Apply through OpenTrain with a free account