I apologize in advance for this crash out, but holy shit I'm so tired. I'm trying to figure out what is even being measured here and it's nearly impossible.
I read the post. It provided no insight whatsoever into how these "tests" work.
Things that were not mentioned:
- Harnesses used
- Tasks used
- How many times tasks are run
- What is done to identify variance in daily runs
- What analysis is done on "bad" runs to detect root causes for failures
- What APIs are being used (matters a LOT)
- How the +/- 10% "variance" was selected
- Why tokens and costs are "weighted" the same, and combined rank roughly as high as intelligence
- Why tokens are included at all when costs are the metric that matters
I'm sorry
@bridgemindai - am I missing something here? I just want to make sure I understand fully before building my own alternative bench.
I DM'd you asking for traces from your runs, would be super helpful as I start digging deeper here.