Freelancer Overview
Advanced Computer Science specialist with extensive experience designing evaluation frameworks, benchmarking complex multi-turn LLM reasoning, and developing custom open-source tools. Proven track record of moving models past superficial optimization by engineering highly structured, atomic, and binary grading metrics. Expert in evaluating high-level programming domains including machine learning architecture, optimization algorithms, and low-level systems execution (Python, C++, ROS 2, Docker).
Key AI Training & Evaluation Expertise:
Evaluation Rubric Architecture (Outlier AI): Experienced in authoring, refining, and executing complex evaluation rubrics to align LLMs on code quality, security, and multi-turn logical consistency. Specialized in breaking down ambiguous prompt criteria into strict, atomic, binary data signals (eliminating subjectivity to ensure high inter-rater reliability) and configuring weighted negative-penalty parameters for structural or logic failures.
Infrastructure & Automation (OpenClaw): Hands-on experience developing and interacting with OpenClaw, demonstrating deep familiarity with handling open-source backend structures, standardizing model outputs, and streamlining workflows for automated and human-in-the-loop data pipelines.
Advanced Code Generation QA: Vetted to review, debug, and optimize complex coding outputs. Expert in evaluating edge cases for multi-threaded code, algorithmic complexity ($O(N)$ optimization), memory management, and specialized ML/robotics scripts.
RLHF & Preference Tuning: Deep understanding of Reinforcement Learning from Human Feedback, prompt engineering, adversarial red-teaming, and generating hyper-specific "hard prompts" designed to stress-test model reasoning.
Core Tech Stack for AI Training
Languages: Python (Advanced/ML frameworks), C++ (Data structures & algorithms), SQL.
Domains: Machine Learning (Inference, LLM security, neural search), Robotics & Systems (ROS 2, Docker, Kubernetes, Linux environment handling).