For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
J
Juan A.

Juan A.

LLM Output Evaluation & Prompt Testing (practicum)

N/A

Key Skills

Software

Other

Top Subject Matter

LLM output quality review and prompt testing
Chatbot/conversational AI QA
RAG vs. direct prompting evaluation

Top Data Types

TextText
DocumentDocument

Top Task Types

Text SummarizationText Summarization

Freelancer Overview

LLM Output Evaluation & Prompt Testing (practicum). Brings 3+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Other, ChatGPT, and Google Docs. AI-training focus includes data types such as Text and Document and labeling workflows including Evaluation, Rating, and Text Summarization.

Labeling Experience

AI-Assisted Research & Documentation (practicum)

DocumentDocumentText SummarizationText Summarization

Organized technical research materials to support AI-assisted documentation workflows. Converted unstructured project notes into structured summaries, readable reports, and clearer workflows. Reviewed and synthesized complex information to improve the usefulness and clarity of outputs for downstream testing or reference. • Structured and summarized research notes into coherent documentation. • Reviewed documentation and produced readable, structured reports. • Synthesized complex information into clearer workflows. • Supported evaluation/testing preparation through organized knowledge outputs.

2024 - Present

RAG vs. Direct Prompting Evaluation (practicum)

OtherTextText

Conducted comparative evaluation between direct prompting and Retrieval-Augmented Generation (RAG) approaches. Designed small test sets, compared model outputs, and verified whether responses were grounded in provided context. Summarized evaluation outcomes in structured tables or short reports. • Built and executed small test sets for model behavior comparison. • Checked grounding in supplied retrieval context for factuality. • Compared response quality differences across prompting strategies. • Produced structured summaries to communicate evaluation results.

2024 - Present

Chatbot & Conversational AI QA (practicum)

TextText

Explored and tested chatbot conversations for quality and reliability across multi-turn interactions. Assessed naturalness, context retention, intent handling, fallback behavior, overconfidence, and recurring conversation failure modes. Documented structured failure patterns and recommended improvement areas based on observed issues. • Reviewed conversational outputs for intent and fallback correctness. • Checked context retention and stability across turns. • Identified overconfidence and inappropriate response behaviors. • Wrote structured notes describing failure patterns and suggestions.

2024 - Present

LLM Output Evaluation & Prompt Testing (practicum)

OtherTextText

Practised evaluating AI-generated responses for accuracy, relevance, clarity, tone, usefulness, and hallucination risk using structured criteria. Compared different prompt variants and documented observations with rubric-style evaluation notes. Produced concise written summaries of findings to support prompt testing and quality improvement. • Evaluated LLM outputs against guideline alignment and quality dimensions. • Performed prompt variant comparison across wording and structure changes. • Used clear evaluation criteria and structured notes for consistency. • Generated short reports summarizing results and failure/strength patterns.

2024 - Present

Work History

S

Self-Directed

RAG vs Direct Prompting Evaluation Assistant

N/A
2024 - Present
S

Self-Directed

Chatbot Quality Assurance Tester

N/A
2024 - Present