Existing AI benchmarks measure if a model completed a task. But at General Legal, we have the data to measure how AI performs against a real human lawyer.
We found that some models can hit close to 90% of the basic requirements. But in practice, they perform much, much worse than what our lawyers actually send clients.
You can read the full post below for all the details.
x.lingyaoai.com/gen_legal_inc/status/2…
We gave 13 AI models legal work to do from our new legal work benchmark. We compared the output of the models with what our human attorneys actually produced. We wanted to know whether a model could complete the whole assignment, not just how many individual requirements it could check off.
Key takeaways for our analysis:
- General Legal's new benchmark evaluates whether AI-drafted legal work is truly client-ready, not just technically complete.
- The benchmark tests agents legal work product with actual attorney work as the gold standard, including access to the firm's proprietary knowledge base.
- Legal work requires judgment about trade-offs, priorities, and strategic choices that extend beyond simply fulfilling a checklist of requirements.
- The benchmark tracks both individual changes and holistic quality, identifying when models add unnecessary terms or make choices attorneys wouldn't release to clients.
- All-pass rate measures the percentage of matters where output meets every criterion, reflecting that even one miss requires attorney intervention.
Read the full article for details.