AI Security Specialist — data labeling, dataset QA, and LLM evaluation for fine-tuning workflows
Evaluated GPT-style outputs for logical consistency, instruction adherence, correctness, bias, and safety while preparing training-ready prompt and response datasets. Built structured instruction-response pairs and preference rankings for fine-tuning small LLMs. Performed Python-based quality checks and flagged labeling inconsistencies to improve dataset reliability. • Developed and executed test prompts for GPT-style models • Conducted rubric-style evaluations of model outputs for factual accuracy and safety • Flagged edge cases and dataset inconsistencies for QA • Designed workflows for AI-assisted cybersecurity analysis and reporting