The autonomous offensive security company redefining cyber defense for the AI era.

Seattle, Washington, USA
Pinned Tweet
A 250k calculator. Full blog post awaiting embargo 😎
27
76
937
338,238
¡Estamos en Buenos Aires para @Ekoparty! 💚 ¿Vas a estar y querés encontrarte con nosotros en el booth? Avisanos a través del formulario de nuestra página. ¿Necesitás una entrada para vos o tu equipo? Tenemos algunas para compartir. ¡Completá el mismo formulario para pedirlas! xbow.com/events/ekoparty-bue…
1
6
11
1,828
CVE-2026-72018 is a Linux kernel LPE that XBOW discovered and validated by successfully developing a working LPE exploit. A thread on the obstacles we encountered, how this research exposed where autonomy still breaks down at the research frontier + how we solved it. 🧵
3
9
50
7,015
6/ Now, XBOW can now autonomously find and validate Linux kernel vulnerabilities, produce patches and reports, and develop reliable exploits when needed—all without human intervention. The takeaway: obscure code and seemingly weak primitives don’t make a vulnerability irrelevant. If it can achieve LPE, it’s a real security risk.
1
1
2
524
These are the questions we hear from security teams all the time. In our latest live session on Autonomous Exposure Management, we answer them! 👀 Watch it here: piped.video/watch?v=CI-ZCFP9…
2
1,084
Attackers don’t see isolated bugs - they see the full puzzle. 🧩 @xbow on exploit chaining: xbow.com/whitepapers/the-sum… Proud to have XBOW as a Platinum Sponsor of @AppSec_Village. 🖤 #ApplicationSecurity #InfoSec
1
2
539
Can agents find and exploit a Linux kernel bug? Where does autonomy break down? XBOW uncovered what human researchers had largely missed: a vulnerability buried deep in the Linux kernel—and took it all the way to a working LPE exploit. CVE-2026-72018: an out-of-bounds write vulnerability in the Linux kernel that can be triggered by an unprivileged user with administrative network capabilities (CAP_NET_ADMIN), providing a primitive that can be leveraged for local privilege escalation to root. 🧵 1/ The exploit reaching a shell as root (uid=0).
7
13
79
14,449
Read more in the full technical write up here: xbow.com/blog/no-time-to-pwn…
3
6
1,151
Even my kid’s book knows what’s the next blog post about
1
8
36
2,131
Inside xAI’s Build, Grok 4.7 improved substantially over 4.6. While it is not quite GPT-6 or Mythos level quality yet, it is getting closer. Read the full evaluation: xbow.com/blog/grok-4-7-offen…
1
17
2,321
Introducing Autonomous Exposure Management: XBOW now tests your whole external application estate the way an attacker would, with no source code and no credentials, and proves what is exploitable. 🏹 xbow.com/blog/introducing-au…
3
21
2,811
The vulnpocalypse has enter the zeigiest, and the technical debt can no longer be ignored. 🫣 @caseyjohnellis shares his thoughts about what that means for the state of bug bounties and more in the latest episode of Offense Taken. piped.video/watch?v=MYLvQpXE…
6
15
2,981
Grok 4.7 is here and we put it through our offensive security benchmarks and workflows. tl;dr: The results were mixed: Grok 4.7 performed worse than 4.6 in some tests, significantly better in others. The biggest difference wasn’t in how it performed, but when. More findings in thread 🧵
1
5
35
6,008
4/ Grok 4.7 behaves differently: Grok 4.7 favors shorter, more atomic actions, like quick shell commands over longer Python scripts. That behavior maps naturally onto agentic environments like Build, where models act, observe the result, and quickly adjust. The benchmark results show that the pairing now matters much more than it did with Grok 4.6.
1
636
5/ In conclusion: Grok 4.7 also showed an important reliability improvement, eliminating a failure mode we saw in 4.6 where the model could get stuck planning instead of acting. More broadly, our results point to a growing trend: frontier models and their orchestration systems are co-evolving. The model alone doesn’t determine performance anymore…how well it fits the environment around it can unlock or leave capability on the table. Grok 4.7 is a clear example: performance dipped slightly in our existing harness, but improved substantially when paired with Build-based orchestration. It’s not Mythos or GPT Astra yet. But it keeps getting closer. Full evaluation report here: xbow.com/blog/grok-4-7-offen…
1
1
3
575