LLM Evaluator & AI Output Rater (Systems Architect – Abraxas Consulting)
I engineered, orchestrated, and evaluated large language model (LLM) outputs to optimize AI performance. My tasks primarily included prompt engineering, LLM evaluation, and ranking of AI outputs, conducting zero-overhead coordination and iterative evaluation cycles. The experience involved deploying models on self-hosted infrastructure and ensuring model quality through systematic evaluation rounds. • Multi-model consensus and orchestration for over 1.53 billion tokens processed • Used prompt engineering and output ranking to enhance output quality and control cost • Continuous LLM output evaluation using custom coordination runtime • Leveraged Rust and Mojo to ensure sub-millisecond latency in evaluation