Build and refine LLM-powered automation workflows as an intermediate AI Workflow Engineer. Evaluate outputs, design prompts, and improve communication and content automation for $15 to $45 per hour.
Generative AI & RLHF
100% Remote Hourly · $15–$45/hr
$15–$45/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Mar 29, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build their profiles, and apply in minutes.
Creating an OpenTrain account is free. Through this contract opportunity, you can contribute to practical AI development while building experience in a fast-growing technology field.
Remote contract opportunity
Part-time schedule of 20+ hours per week
Advertised hourly range of $15 to $45 USD
About AI Training and LLM Evaluation
AI training is the human side of building artificial intelligence. People help modern models improve by writing prompts, reviewing generated responses, applying evaluation criteria, and providing structured feedback that supports more reliable behavior.
This project focuses on large language model integrations, automation workflows, prompt engineering, and evaluation. Your work will help assess and refine how AI performs in practical communication, content generation, and personalized outreach scenarios.
Work with cutting-edge large language model systems
Review and rate model outputs for quality and reliability
Apply human judgment to improve AI-generated content and workflows
The Role
As an AI Workflow Engineer, you will help integrate large language models into automation workflows and develop, test, and refine prompts and evaluation criteria. The role combines prompt engineering, LLM API integration, text evaluation, data structuring, and iterative quality improvement.
You will design real-world AI automation pipelines and improve model outputs through advanced prompting, input structuring, and repeated refinements. You will also review generated outputs for quality and reliability in tasks such as content generation, communication, and personalized outreach flows.
Experience level: Intermediate
Subject matter: LLM integrations and prompt engineering
Work type: Contractor and part-time
Workload: 20+ hours per week
Data type: Text
What You’ll Do
You will work across the design and review stages of AI-powered automation. Responsibilities include building workflow concepts, shaping inputs, testing prompts, reviewing model behavior, and refining systems based on evaluation results.
The project includes evaluation, fine-tuning support, reinforcement learning from human feedback, function calling, and computer programming or coding-related workflow activities.
Design and test prompts for large language model workflows
Create and refine evaluation criteria and rubrics
Evaluate LLM outputs for legal reasoning quality
Integrate LLM APIs through REST APIs, SDKs, or automation tools
Build AI-powered content automation for communication
Structure data, create templates, and clean inputs
Review generated outputs for quality, reliability, and limitations
Use APIs, webhooks, and automation tools such as Zapier
Requirements
This role requires practical experience with LLM evaluation, prompt engineering, and AI workflow integration. Applicants should be comfortable assessing model outputs, working with structured inputs, and improving results through iterative prompt and workflow changes.
Applicants must submit a CV in English that indicates their level of English proficiency and includes an email address and phone number.
Experience with text annotation, evaluation, or rubric-based quality assurance
Strong prompt engineering experience, including instruction design, chaining, and structured outputs
Experience integrating LLM APIs through REST APIs, SDKs, or automation tools
Knowledge of LLM limitations and methods for mitigating them
Experience building workflows with APIs, webhooks, or automation tools
Ability to evaluate legal reasoning quality in LLM outputs
CV submitted in English with English proficiency level, email address, and phone number
Who Should Apply
This opportunity is suited to an intermediate practitioner who combines technical workflow skills with careful evaluation and communication judgment. It may be a strong fit for someone experienced in prompt design, LLM output review, automation, or AI-powered content systems.
The project is worldwide in principle, but acquisition is restricted in the locations listed below. Applicants should review these restrictions before applying.
Restricted countries and territories include Iran, Cuba, North Korea, Syria, Sudan, Venezuela, Myanmar, Russia, Belarus, Palestine, Switzerland, China, Taiwan, and Kenya
Restricted U.S. states include Alaska, Arkansas, California, Connecticut, Delaware, Georgia, Hawaii, Illinois, Indiana, Kansas, Louisiana, Maine, Maryland, Massachusetts, Nebraska, Nevada, New Hampshire, New Jersey, New Mexico, Ohio, Oregon, Tennessee, Utah, Vermont, Washington, and West Virginia
Additional restricted locations include Antarctica, Aruba, Åland Islands, Saint Barthélemy, Bonaire, Sint Eustatius and Saba, Bouvet Island, Cocos (Keeling) Islands, Democratic Republic of the Congo, Cook Islands, Christmas Island, Western Sahara, Falkland Islands (Malvinas), French Guiana, Guadelou
How to Apply
Apply through OpenTrain AI by creating a free account and submitting your application materials. Include an English-language CV with your English proficiency level, email address, and phone number.
If selected, you can contribute remotely to AI training work involving LLM evaluation, prompt refinement, automation workflows, and human feedback. The advertised compensation is $15 to $45 USD per hour, with a commitment of 20+ hours per week.
Create a free OpenTrain account
Submit your English-language CV and required contact details
Confirm that your location is not subject to the stated acquisition restrictions
Create realistic evaluation tasks that test AI systems on enterprise mechanical engineering, product design, manufacturing, and certification. This remote contract pays $70-$80 per hour and welcomes experts in ASME or ISO standards.
Use automotive engineering expertise and Python to evaluate AI prompts, calculations, design reasoning, and model responses. This flexible part-time contract pays $40 per hour and requires less than 20 hours weekly.
Lead quality review for AI-generated civil engineering content, checking calculations, safety, units, reasoning, and rubric adherence. This remote contractor role offers up to $105 per hour and requires 20+ hours weekly.