Linguistic Data Refinement (Independent AI Training)
Conducted independent AI training by comparing LLM outputs and labeling responses for factual truthfulness, helpfulness, and harmlessness. Investigated and flagged edge cases where models missed cultural nuances or local context, with emphasis on West African linguistic structures. Authored gold-standard responses used to calibrate automated evaluation scripts. • Labeled model responses along RLHF criteria: truthfulness, helpfulness, and harmlessness • Identified failure modes and documented edge cases related to culture and local context • Produced gold-standard reference answers for evaluation calibration • Supported evaluation workflows with structured, logically consistent labeling outputs