This is a message I shared with the team earlier today. It's still day one at Crosby.
21
14
405
51,775
John Sarihan retweeted
We love a game of strategy and execution. Crosby is hosting a VIP box at League of Legends Worlds Finals in NYC for members of the community, recovering and current league addicts alike. Apply below.
31
13
118
36,530
John Sarihan retweeted
At @crosbylegal we see a future where you aren't nervous to email your lawyer. How many $1000s will that cost? When will I hear back? That future is already here. LLMs are the greatest thing thats ever happened to our profession. Excellent legal work can be delivered at a fraction of the old speed and cost. So how is it possible that law firms are still billing by the hour? Our profession can do better. It must.
2
8
27
1,944
John Sarihan retweeted
A hearse. From Lux family co @crosbylegal
RIP the billable hour. It will be remembered fondly by everyone who charged it, and by no one who paid it. In lieu of flowers, send contracts to crosby.ai
3
3
23
9,396
John Sarihan retweeted
CROSBY CARS IN MY HOOD
3
3
34
3,805
RIP the billable hour. It will be remembered fondly by everyone who charged it, and by no one who paid it. In lieu of flowers, send contracts to crosby.ai
8
6
88
31,292
John Sarihan retweeted
In Memoriam: The Billable Hour, 1913-2026. Procession hours: Tues 9/15 12-8 pm Wed 9/16 8 am-6 pm San Francisco, CA
13
7
78
16,308
John Sarihan retweeted
‘@crosbylegal for X’ is the new ‘Uber for X’
4
2
39
23,288
John Sarihan retweeted
i love when my friends' products work 1/ @crosbylegal gets me through brutal redlines everyday with 0 law degree 2/ @trymelius made me a sick concept video i pitched to the big boss and it landed 3/ @bot manages my complex data evals and produced me a rap lyric video 4/ @tryramp router saves me a ton of $$ and implemented my feedback on the same day SO PROUD
1
2
53
9,552
What does it take to align not just AI systems, but the institutions and laws that govern them? On September 1st, Crosby Intelligence is hosting @PeterHndrsn for a fireside chat at the Crosby office. Peter Henderson is a professor at Princeton with a Stanford JD and PhD in computer science. After working at Meta FAIR, his research now spans reinforcement learning, legal alignment, human <> AI collaboration, and the choices the US may face in a post-AGI world. Crosby researcher and Stanford CS PhD candidate Teddi Worledge will lead the conversation, covering: - RL for long-horizon decision-making - Human <> AI collaboration in law - Why legal alignment is the next major frontier We’re bringing together researchers from across academia and industry. If you work on RL, AI alignment, governance, we’d love to have you in the room. Space is limited- sign up here: partiful.com/e/CItDsRE6n0q9O…
3
3
21
3,600
We don’t talk much publicly about what we’re building internally. Our main goal is for lawyers to review an agent’s work, rather than redline themselves. The better the agent is, the less time a lawyer spends reviewing its work. Last quarter we finally hit a longtime goal: “the 5 minute NDA review.” We’re the first company to reliably achieve this. Here's how we solved a category of difficult NDAs. Sharing a slide from our last board deck:
Crosby (@crosbylegal) is a “neo-firm," equal parts lawyers and engineers that owns liability, and takes on legal work end to end, handling contracts for @ramp, @clay, @cursor, @cognition, @ListenLabs and more. I spent two weeks embedded in Crosby in New York - seeing their time trials (internal contests where lawyers race to redline a contract), All Hands, eng meetings, hung out in their Slack. I even shadowed lawyers. Over 25 conversations later, I'm excited to share this field study going deep on solving unverifiable problems in law, profile on @ryanjdaniels (Stanford-trained attorney) & @jsarihan (early employee at Ramp), and how "misfit lawyers" feel about the role of the lawyer changing. Read it here - new-ontologies.com/posts/Cro…
3
7
60
7,546
John Sarihan retweeted
Often we just hear about companies when they're announcing milestones, so it's always great when writers take time to embed themselves in an organization and really see how they work. Nicole did that with @ryanjdaniels, @jsarihan, and @crosbylegal recently. If you're interested in what it looks like inside an early stage AI startup, especially one that has actual lawyers who are dogfooding their own product, this is worth a read. "No one knows what legal work becomes in the long term, and I wouldn’t trust anyone who spouts oracle wisdom. But inside the neofirm I saw glimpses of the old world vanishing, and hierarchical systems being flipped. If lawyers can spend less time on deadweight tasks, their time becomes even more relational - building mutual context and trust between strangers, where the process of getting to an agreement is an end in itself. The lawyer remains more than the last mile. After all, “a contract review is a conversation between two humans,” Ryan said as we walked up and down Crosby street. “It entails feeling.”"
Crosby (@crosbylegal) is a “neo-firm," equal parts lawyers and engineers that owns liability, and takes on legal work end to end, handling contracts for @ramp, @clay, @cursor, @cognition, @ListenLabs and more. I spent two weeks embedded in Crosby in New York - seeing their time trials (internal contests where lawyers race to redline a contract), All Hands, eng meetings, hung out in their Slack. I even shadowed lawyers. Over 25 conversations later, I'm excited to share this field study going deep on solving unverifiable problems in law, profile on @ryanjdaniels (Stanford-trained attorney) & @jsarihan (early employee at Ramp), and how "misfit lawyers" feel about the role of the lawyer changing. Read it here - new-ontologies.com/posts/Cro…
3
7
73
26,001
John Sarihan retweeted
Crosby (@crosbylegal) is a “neo-firm," equal parts lawyers and engineers that owns liability, and takes on legal work end to end, handling contracts for @ramp, @clay, @cursor, @cognition, @ListenLabs and more. I spent two weeks embedded in Crosby in New York - seeing their time trials (internal contests where lawyers race to redline a contract), All Hands, eng meetings, hung out in their Slack. I even shadowed lawyers. Over 25 conversations later, I'm excited to share this field study going deep on solving unverifiable problems in law, profile on @ryanjdaniels (Stanford-trained attorney) & @jsarihan (early employee at Ramp), and how "misfit lawyers" feel about the role of the lawyer changing. Read it here - new-ontologies.com/posts/Cro…
37
21
263
125,971
There’s something very special about the way people’s paths diverge and then cross again. Nicole and I knew each other in college, long before Ramp or Crosby. Years later, she spent two weeks embedded with our team and wrote this thoughtful, deeply observed piece about what we’re building. A little surreal, and very meaningful, for her to write such an incredible piece that snapshots this chapter at Crosby so well.
Crosby (@crosbylegal) is a “neo-firm," equal parts lawyers and engineers that owns liability, and takes on legal work end to end, handling contracts for @ramp, @clay, @cursor, @cognition, @ListenLabs and more. I spent two weeks embedded in Crosby in New York - seeing their time trials (internal contests where lawyers race to redline a contract), All Hands, eng meetings, hung out in their Slack. I even shadowed lawyers. Over 25 conversations later, I'm excited to share this field study going deep on solving unverifiable problems in law, profile on @ryanjdaniels (Stanford-trained attorney) & @jsarihan (early employee at Ramp), and how "misfit lawyers" feel about the role of the lawyer changing. Read it here - new-ontologies.com/posts/Cro…
4
12
66
13,233
John Sarihan retweeted
Crosby has been named to @Forbes ' Next Billion Dollar Startups list alongside some of our favorite customers like @gumloop & @pacecom (who wisely chose today to onboard to Crosby). We are grateful to be recognized alongside companies we deeply respect like @juicebox_work, @fortellresearch, @Netic_AI, @cogent_security, and @ValthosTech. Our team has worked hard over the past year to deliver high quality legal work to some of the fastest growing companies in the world, including @cursor_ai, @cognition, and @tryramp. We have an ambitious roadmap and are hiring across engineering, research, legal, and many more roles.
3
4
55
58,106
John Sarihan retweeted
We're looking for the outliers. The lawyers who love their work, their clients, the law... but not their law firm, the formality, the slowness, the aversion to change. The lawyers who know there is a better future begging to be built – without billable hours or extortionate fees – and who are dying to build it. We're a team of misfit attorneys and we're having the time of our lives. Come join us. crosby.ai/lateral
3
7
38
7,172
Benchmarks are everything! As evals make more tasks measurable, the difference between a $1,000/hour and a $3,000/hour partner will be a proven benchmark metric.
The verdict is in! Frontier models can pass the bar, yet they struggle on comprehensive legal research Today we're releasing Legal Research Bench, a benchmark that measures models’ ability to solve realistic legal research tasks across eight areas of U.S. law Instead of awarding partial credit, Legal Research Bench measures whether a model can conduct exhaustive legal analysis. We grade against a strict, all-pass rubric written by practicing lawyers. A model only receives full credit if every required legal element is correct Claude Opus 4.8 leads with 43.8% all-pass accuracy, followed by GPT 5.5 (40.4%) and Claude Sonnet 4.6 (38.5%). While top models score around 80% with partial credit, none exceed 44% when every required legal element must be correct The gap between partial and all-pass accuracy shows how difficult it remains for AI to produce complete, reliable legal research. We hope that Legal Research Bench helps better measure, and ultimately close that gap Lots of exciting work happening in Legal AI from @harvey and @crosbylegal. Excited for the legal research benchmarks ahead!
39
8,458
John Sarihan retweeted
POV you're unprepared for a podcast with openai's two newest researchers @tbpn @jordihays @johncoogan
4
71
8,780
John Sarihan retweeted
"It's not the same as a math problem where you can just verify it." @johncoogan @jordihays & @ryanjdaniels break down benchmarks for non-verifiable domains and RedlineBench on @TBPN
1
2
14
2,153
John Sarihan retweeted
1. as a mental model it is more correct to think of fable+ class models as english -> code interpreters - converts your idea into code into "correct" code regardless of problem complexity and output complexity (diff size). Fable 5 will be the worst of this new class of models 2. diff size/complexity is to be managed purely for review: small diffs - in high risk areas of code (auth/identity/data access/network access/money movement) large diffs for code that can be empirically verified (frontend/backend plumbing/code without network or db access/performance code that can be empirically verified) 3. time it takes to ship software is completely disconnected from time to produce the PR - how long the work takes depends fully on ability to review/merge code while managing risk at scale 4. solving the bottlenecks for above matter enormously- linters/testing/CI/shadow mode verification/empirical verification 5. agency matters enormously- what are the biggest bottlenecks to speeding up the loop and eliminating them? what are the problems that need solving and when do they need solving? what does it take to the solution to all of them today? 6. deep understanding of the full stack matters enormously- what problems are worth pursuing? is there a higher level of problem abstraction to address first? should I give it the sub-sub task, the sub task, or the task itself. what are the major risks with this PR (order of importance: security holes/correctness holes/performance holes). is there a higher speed way of producing data that allows me to merge this? should this be run in shadow or in a sandbox or a flag. understanding every line of logic may not be needed but understanding and managing risk matters enormously. 7. the cost of complexity itself is changing. it might be now worth "maintaining" 50% more code to get a 5% performance win. getting the right abstractions matter less because larger refactors are less tedious. code quality nits become huge drag. very likely, a much smarter model will be maintaining your code so worth taking on more technical debt now. taking the time to hand architect and rebuild systems comes with an enormous cost of velocity 8. if it quacks like a duck and walks like a duck, it's a duck. For low risk cases, it might be more sane to treat code chunks (services / functions) as a black box, like we do for neural networks: do full empirical verification only: has code produced correct outputs for the last 10,100,1000,10k inputs ? can we quarantine this large piece of code - no outbound access to network / database ? what happens when this code is wrong? do we get hacked/or crash(memory/cpu)/is an inconvenience? is it internal facing or external? what can we do to address these risks? 9. eventually, logical verification (line by line review) will come at an enormous cost- save it for where it matters and build systems that are tolerant to empirical verification. is there a decorator that prevents db / network access? correctness bugs are significantly easier to rectify than access bugs 10. what are the rails that allow for even faster iteration? code permissions can be opt in - db writes, db reads, network egress (to where?), PII access. how long does it take to get shadow mode data? how many PRs can be tested? What are the categories of diffs
66
148
1,777
311,724