Retrieval-Augmented Generation, Explained Properly
The idea is simple and the failure modes are not. Most disappointing implementations fail at retrieval, not at generation.
The idea is simple and the failure modes are not. Most disappointing implementations fail at retrieval, not at generation.
You cannot unit-test a language model, which is not the same as being unable to test it. A small hand-built evaluation set is worth more than any public benchmark.
Context windows have grown enormously. The ability to use everything in them has not grown at the same rate, and the gap matters.
Confident invention is a consequence of what these systems are trained to do. It can be reduced substantially and it cannot be eliminated by prompting.
Arguing about a hypothetical future system is more comfortable than arguing about the deployed ones. That is largely why it is so popular.
A number that improves every quarter is reassuring and largely decorative. Measuring what a system is good at is much harder than ranking it.