$ cat /etc/cookies.conf
We use cookies to understand how people use this site.
Analytics cookies help us improve your experience.
They are off by default. Nothing tracks you until you say so.
$ select cookie_preferences
Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Learn how to build reliable LLM-based evaluations using only about twenty human annotations, achieving scalable, human‑aligned assessments that outperform existing benchmarks.
Evaluating AI apps is often only possible by humans reviewing outputs or using LLMs to evaluate them. The former is laborious, slow, and expensive. The latter often falls short of correctly evaluating the outputs.
We’ve developed a method to create reliable LLM-based evals with as few as 20 manual annotations. Allowing you to automagically turn “vibe checks” into scalable and reliable evaluations aligned with human judgment.
We’re also releasing preliminary results on JudgeBench on which our automatically created LLM eval using gpt-3.5-turbo beats gpt-4o.
Loading recent emails...