AI agent evaluation is becoming a real bottleneck as more teams ship to production.
Most ICs (myself included) don't enjoy writing tests. We've built in-house eval pipelines but haven't tried third-party platforms yet.
How's your team handling this — in-house or outsourced?👇