LLM Evaluation Software Engineer Ruby
Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.
Posted Jul 20, 2026
Build Java backend components and evaluate large language model responses for relevance, clarity, accuracy, and safety. This worldwide, part-time contractor role offers 20+ hours per week for engineers interested in shaping AI.
Coding & Software
Worldwide
Eligibility
Entry
Experience
Jul 17, 2026
Posted
Open worldwide
OpenTrain AI is the hiring and contracting organization for this role and the #1 platform for finding and building careers in AI training and data labeling. OpenTrain helps contributors build a lasting profile, discover opportunities across the industry, and grow their experience in a rapidly developing field.
AI training is the human side of building modern artificial intelligence. Engineers, annotators, and subject-matter experts prepare examples, review model outputs, and provide feedback that helps AI systems become more useful, accurate, and aligned with user needs.
In this role, your software engineering knowledge will support evaluation of dialog agents and large language models. Your judgments, technical explanations, datasets, and code contributions can help improve systems used for education, entertainment, and general question answering.
OpenTrain AI is seeking a Java LLM Evaluation Engineer to develop and maintain high-quality backend code for AI model training and optimization. The role combines practical Java development with hands-on evaluation of dialog agent systems.
You will assess model responses against defined criteria, explain your decisions, create task-specific training data, and help improve evaluation strategies. You will also contribute to supervised fine-tuning and reinforcement learning from human feedback, including work related to reward model refinement.
You will combine backend engineering, software quality practices, and structured AI evaluation. Clear technical reasoning and careful attention to evaluation criteria will be important throughout the work.
You should be comfortable writing practical Java and working with backend development concepts. The role requires strong English communication and the ability to follow detailed evaluation criteria while documenting technical judgments clearly.
A bachelor's or master's degree in engineering or computer science, or equivalent experience, is appropriate. Prior software quality assurance or test planning experience and experience evaluating model responses or creating training data are helpful.
This opportunity may suit an entry-level software engineer, Java developer, QA professional, or technically minded AI contributor who wants to apply programming and evaluation skills to real model-improvement work. You do not need a separate AI-training career history if you can demonstrate the required Java, backend, communication, and analytical abilities.
Create a free OpenTrain account and apply in minutes. Your OpenTrain profile can help you present relevant software, evaluation, and AI-training experience as you discover and pursue future opportunities in the field.
Keep exploring
Build and evaluate real-world Ruby software engineering tasks for LLM training datasets. This remote contractor role offers 20, 30, or 40 hours weekly with required PST overlap.
Posted Jul 20, 2026
Help train and benchmark large language models by writing, correcting, and evaluating production-quality code across multiple languages. This flexible, worldwide contractor role requires 20+ hours weekly and is available through OpenTrain.
Posted Jul 16, 2026
Build and evaluate challenging C++ software engineering tasks that help measure how well large language models understand and fix real code. Work remotely for 20 or more hours weekly through OpenTrain.
Posted Jul 17, 2026
Browse related job pages
Languages