LLM-as-judge and evaluation rubrics
LLM as judge uses a capable language model to score outputs against a rubric. It scales evaluation to thousands of examples at a fraction of human evaluation cost. The trade off is bias: judge models favor their own outputs, longer responses, and confident sounding but incorrect answers.