Independent AI Developer and Researcher — Ferrum Flux Fenice
Conducted an LLM behavioral study to evaluate model stability across prompt-based workflow design. Tested six open-source large language models at multiple temperature settings using a structured dataset and recorded results in JSON. Iteratively refined evaluation criteria to distinguish reliable versus unstable behaviors for downstream automation use. • Designed evaluation criteria for model behavior • Executed testing/validation on a 1,440-row structured dataset • Documented a 144,000-inference study in full JSON • Applied iterative improvements based on findings for prompt-based workflow reliability