Evaluating an AI Feature Before You Ship It
You cannot unit-test a language model, which is not the same as being unable to test it. A small hand-built evaluation set is worth more than any public benchmark.
You cannot unit-test a language model, which is not the same as being unable to test it. A small hand-built evaluation set is worth more than any public benchmark.
A number that improves every quarter is reassuring and largely decorative. Measuring what a system is good at is much harder than ranking it.