Prompt Engineer & AI Response Evaluator (Luel)
Crafted targeted prompts to test AI reasoning, factual accuracy, and instruction-following behavior. Evaluated AI outputs across multiple dimensions including tone, relevance, safety, and factual correctness. Flagged hallucinations, biases, and unsafe content using structured feedback and detailed annotation notes, including for multilingual evaluation. • Prompt generation for targeted capability testing • Multidimensional output evaluation (safety, relevance, correctness) • Hallucination, bias, and unsafe-output annotation • Multilingual evaluation and performance assessment across languages