LLM post training intern, Ethara AI (Remote)
Performed supervised post-training for large language models to improve response quality, instruction following, and domain performance. Created and reviewed annotation datasets used for preference ranking, response evaluation, and safety/alignment labeling. Evaluated model outputs for accuracy, coherence, factual consistency, and policy compliance to support iterative improvement workflows. • Preference ranking label creation and dataset curation. • Response evaluation and safety/alignment labeling dataset work. • LLM output evaluation for factual consistency and policy compliance. • Contribution to benchmarking and iterative model improvement cycles.