Yesterday I hosted a fun dinner conversation with
@vincentsunnchen from
@SnorkelAI on evals and RL environments.
The "data and RL env" companies (like Snorkel) have seen massive growth in the past few years. There's been an explosion of interest in evals. At the same time, models are ripping through benchmarks with each new release. We talked about the evals everyone is defining, what evals are still left unsolved, what’s left up to frontier models vs. intelligence that you own, and more:
- A big challenge for building RL environments is “fairness” - when the model fails on a given environment, can you attribute it to the input, harness, or reward model?
- Building proper rewards is hard. Some tasks are not easily quantifiable. You also want to discourage reward hacking. At the same time, you don’t want to be too prescriptive with intermediate rewards.
- Long horizon evals are still extremely hard, some business processes can take up to weeks or months before the final outcome
- Most regulated industries still need human in the loop to guarantee ~100% accuracy, “80%” accuracy is not good enough
- As models get more intelligent, there will be a barbell of boutique data vendors (e.g. any SMB) any scaled up data providers.
- Models still exhibit “jagged intelligence” where they still fail on a long tail of edge cases.
- There might always be opportunities to gather unique data for a given task and posttrain models for lower cost and higher accuracy.
This marks #003 in our founder dinner series. What topic should we discuss next? Let us know your thoughts below!