Build automated data-quality tests and evaluate PII and PHI de-identification systems for privacy leakage, accuracy, and reliability. This fully remote contract requires strong Python, SQL, ML evaluation, and privacy expertise.
About OpenTrain
OpenTrain AI is hiring and contracting for this remote role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping professionals discover projects, build a credible portfolio, and grow in a fast-moving technology field.
- Fully remote AI training and engineering work
- Create a free OpenTrain account and apply in minutes
- Build a profile that showcases your AI-related experience
About AI Training and Data Quality Work
AI systems depend on carefully prepared, tested, and reviewed data. Engineers and evaluators help make these systems more reliable by validating datasets, measuring model behavior, identifying failures, and improving the safeguards around sensitive information.
- Work on cutting-edge machine learning and data reliability problems
- Use automated testing and human review to improve AI systems
- Contribute remotely with flexible contract opportunities across the industry
The Role
OpenTrain is recruiting an ML and Data Engineer focused on data quality, machine learning evaluation, and privacy compliance. You will improve the reliability of enterprise data pipelines and test PII and PHI de-identification systems so sanitized data remains accurate, consistent, and free of sensitive information leaks.
The role combines automated testing, adversarial dataset design, quality analysis, and compliance-focused review for AI systems. The work is fully remote and structured as a contract engagement.
- Role focus: PII de-identification system evaluation
- Contract type: Contractor and part-time
- Initial contract duration: One month, with possible adjustment based on the engagement
- Compensation: Not disclosed
What You'll Do
You will design evaluations and quality controls that identify defects across enterprise data pipelines and machine learning systems. Your work will span data validation, adversarial testing, privacy leakage analysis, monitoring, and compliance documentation.
- Design validation suites for schema integrity, completeness, drift detection, and reconciliation across pipeline stages
- Analyze enterprise data sources and connectors for coherence, domain coverage, consistency, and depth
- Build adversarial test datasets covering obfuscated identifiers, multilingual entities, OCR noise, edge cases, and unusual document formats
- Evaluate NER and ML-based de-identification systems using precision, recall, F1 score, leak rates, and false-negative analysis
- Investigate privacy leakage across raw, processed, and sanitized data
- Integrate regression gates, monitoring, reporting, sampling audits, and compliance trails into data and engineering workflows
Requirements and Technical Skills
This role requires substantial hands-on experience in data engineering, ML engineering, data quality, or a related field. The structured role information identifies the opportunity as entry level, while the role requirements call for five or more years of relevant experience; applicants should review that experience expectation carefully.
- Five or more years of experience in data engineering, ML engineering, data quality, or a related field
- Strong Python and SQL skills for test automation, data validation, and ML evaluation
- Knowledge of PII, PHI, de-identification concepts, and privacy frameworks such as HIPAA Safe Harbor, GDPR, or LGPD
- Experience evaluating NER or other non-deterministic ML systems with labeled datasets
- Ability to use precision, recall, F1, leak-rate, and false-negative metrics
- Ability to design reliable evaluations for non-deterministic ML or LLM-based systems
- Experience integrating automated tests and quality checks into CI/CD pipelines
- Familiarity with pytest, Great Expectations, Pandera, BigQuery, Google Cloud Storage, or Cloud Run Jobs is valuable
Schedule and Remote Contract Details
The role is fully remote and requires 40 hours per week, including six hours of daily overlap with Pacific Time. The structured opportunity information also lists a time requirement of 20 or more hours per week, so confirm the expected commitment during the application process.
- Worldwide remote opportunity
- English-language role
- Required overlap: Six hours daily with Pacific Time
- Stated schedule: 40 hours per week
- Structured time requirement: 20 or more hours per week
Build Your AI Engineering Career With OpenTrain
OpenTrain gives AI professionals one place to manage opportunities, present credible experience, and develop a lasting portfolio in AI training and data labeling. Create a free account to build your profile and apply for work that matches your technical background.
- Showcase data engineering, machine learning, and evaluation experience
- Discover specialized AI work in one place
- Turn project experience into a stronger long-term professional portfolio