Tool-use evaluation, Tool-augmented LLM evaluation, RAG, Multimodal comprehension Test | Micro1 & Mercor
Evaluated whether the model can generate effective search queries and correctly use external sources under retrieval constraints. Performed retrieval-augmented generation (RAG-style) and multimodal comprehension evaluation using video understanding via transcripts or captions. Assessed the model’s ability to retrieve relevant information with search tools and produce accurate instruction-following summaries. • Tested tool use for search query generation • Verified retrieval constraint adherence and source usage • Evaluated multimodal/video comprehension from transcripts or captions • Rated summary correctness for instruction-following outputs