Skip to main content

3 docs tagged with "evaluation"

View all tags

Eval Datasets & Synthetic Data

How to build golden eval datasets for LLM systems, label them consistently, split them safely, and use synthetic data without fooling yourself.

Evaluation & LLMOps

How to test non-deterministic LLM systems with datasets, scorers, and LLM-as-judge; eval-driven development and harness engineering; and the LLMOps discipline of operating prompts, models, and agents in production.