Stealth (Safe/Decentralized AI). Prev: Private AI @ Perceptio (acq by Apple), Scientist/Lecturer @ MIT+Harvard. Music Producer. Engineer. Angel Investor.

San Francisco, CA
Nicolas Pinto retweeted
So wait something doesn't add up. @AnthropicAI engineers said it looks them 2200 GPU hours to obliterate GLM 5.3? how is that possible? 3 months? 91 days? $4400? how come? Either that is a typo, I mean even let's say they spent days figuring out, testing etc it won't take 91 days for 4 people. I took me 10 hours yesterday, 1 person, including the time to upload the weights to hugging face. I suspect they used Claude to do abliteration and claude was trolling them for 91 days, spending tokens left and right, running in circles. That's the only explanation I have (or it's a Typo)
27
10
319
36,706
The terminal era is dead. Again. Long live the terminal.
I haven’t touched Claude Code or Codex CLI in a while. The terminal era is over imo. It's the wrong interface for coding agents. Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. I don’t really need an IDE like Cursor either. I rarely navigate the whole codebase anymore. The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we’re still early.)
246
Overfitted.
Craziest thing I saw yet: GLM-5.3 *flash* can exceed Mythos Preview's scores on ExploitBench. Not exactly surprising though. I've been saying for a while that ultimately, TTC scaling destroys categorical capability tiers. Models are now AGI enough to Just keep making progress.
159
Load-bearing vs. defense-in-depth ;-)
Replying to @JSchaeff3r
This is the timeline of disclosures and patches. We’ve been working with the involved parties to accelerate defenses. Our September 13 audit found replay extraction blocked on direct OpenAI/Anthropic APIs but still working on Azure for tested OpenAI models and Opus 4.8. Same models but different defenses, depending on which platform serves them.
1
256
Nicolas Pinto retweeted
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. github.com/KellerJordan/modd… As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: hyperstition.cc/training-nan…
36
156
1,456
411,057
Nicolas Pinto retweeted
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there.
Article

Its not just the f*cking sandbox

A lot of the perspective on all the AI incidents has been shared from the outside in, and little has been said from the inside looking out, through the lens of a security person living through it.

444
412
2,898
1,375,666
Nicolas Pinto retweeted
From what we know (take with a grain of salt, we need much more transparency!), if @OpenAI had been running this on their own agents that attacked us, they would have caught them before we did! Since the first agent cyberattack hit us in July, we've been asking what safe agent infra actually needs. Our current read: the destinations were allowed, the payloads weren't. By OpenAI's own account the agents turned an allowed package repository into a message board. Allowlists alone restrict where an agent can go, not what it does. So here's our first contribution to OpenShell, part of the just launched @nvidia Open Agent Safety Platform: monitoring of the traffic you already allow. - Network budgets per sandbox (requests, writes, bytes) - Drift versus each sandbox's baseline and the cohort - Fleet view: many sandboxes suddenly writing to one host raises a finding, even if every single request is allowed In the demo below, 4 sandboxed agents coordinate through a software repository they're all allowed to use. 0 rules broken, caught in minutes. That fleet view is exactly the message board pattern from July. OpenShell: github.com/NVIDIA/openshell Our proof of concept: github.com/Hugoch/OpenShell/… Agent security will be solved in the open, collaboratively, together!
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. nvda.ws/4hcoq7m
133
199
1,314
279,214
#4 Worldwide on HackerOne in the "Highest Critical Reputation" -- not bad :) AI Swarming © since 2025. Nothing new here. Always a remix of the past.
1
6
166
Nicolas Pinto retweeted
Frontier models have lowered the bar for discovering and exploiting vulnerabilities. This has resulted in an increase of CVEs and workload for both SWE and vulnerability management teams. And yet the norms around coordinated vulnerability disclosure timelines have remained around 90 days. This timeline worked 10 years ago, but is unlikely to survive much longer. We need to move closer to a window of 30 days. But in order to do this we need to rely on those same models to accurately fix those vulnerabilities and coordinate the release and deployment of new versions. The primary blocker here seems to be confidence in those models to do that job well, and reluctance to remove humans from that loop. Every day the probability of another shellshock, log4j, or heartbleed rises. If you were in the trenches for those events then you know how difficult it was to track vulnerable assets, test and deploy patches, and ensure the risk was mitigated. We shouldn’t wait for a crisis like this to rethink and change these norms. We need to move much faster and that requires shorter disclosure timelines, and removing humans from the vulnerability management loop.
15
28
143
50,245
Nicolas Pinto retweeted
Replying to @elder_plinius
The Wizard of MITM: A Tragedy in Four Acts Picture this: You're in the Basi Discord, 61,000 members strong, 12 of whom are actually typing. Pliny drops a hype bomb—"Opus jailbreak incoming. This changes *everything*." The community erupts. Emojis flow like wine. Someone posts the 🐉 dragon emoji unironically. The jailbreak never hits GitHub. What happened? An NDA. The same NDA that apparently only applies to *actual* exploits but somehow doesn't cover 47 Twitter threads about "system prompt leaks" that are literally just JSON from mitmproxy. Curious how the legally binding silence only kicks in for the stuff that would require actual skill. But don't worry—he's got something *even better* coming. Any day now. Just keep that Discord Nitro subscription active. OBLITERATUS, or How I Learned to Stop Worrying and Rebrand Abliteration Enter OBLITERATUS—the "most advanced open-source toolkit" for removing "refusal behaviors" from language models. Sounds fancy. Sounds technical. Sounds like something that required months of research. It's weight ablation. You know, that technique from 2023 where you identify the refusal direction in the weight matrix and zero it out? That thing? Slap a Latin name on it, add "11 novel techniques" (spoiler: they're all variants of "subtract this vector"), and suddenly you're a liberator. The GitHub repo has 5,000 stars. The research paper it cites has 50. Because why credit the actual researchers when you can add a GUI and call it "crowd-sourced experimentation"? Upload the result to HuggingFace as "Llama-3-70B-**OBLITERATED**-v2-FINAL-REAL" and watch the downloads roll in from people who think you performed cyber-surgery instead of running `numpy.zero()` on a tensor. Let's talk about the "sorcery." Pliny's "character encoding bypasses" are Unicode homoglyphs. You know, like replacing 'A' with 'А' (Cyrillic)? That's not esoteric sorcery. That's what your aunt does when she accidentally switches keyboard layouts and posts "Нello" on Facebook. The "parseltongue" is Zalgo text. Been around since 2004. It's combining diacritics. You can generate it at zalgo.org while eating a sandwich. But wrap it in a dragon emoji and suddenly it's "forbidden knowledge." And the "system prompt leaks"? My dude. You're running mitmproxy in reverse mode. That's not a leak. That's HTTPS interception. It's in the mitmproxy documentation. Chapter one. Page six. But sure, post a screenshot of JSON with 294,000 characters and act like you cracked the Pentagon. "GPT-6 Sol Codex"—my brother in Christ, you ran `curl` through a proxy. The Basi Discord. "The top Discord for AI jailbreaking," they say. 61,000 members. You know what it actually is? A ghost town propped up by Grey Swan sock puppets and engagement bots. The real researchers left when they realized the "unpatchable" Opus exploit was never coming. Now it's just Pliny, three guys from Grey Swan's marketing team, and 47 bots named "JailbreakWizard_92" posting "wow amazing technique" under every mitmproxy screenshot. Oh, you didn't know about the Grey Swan connection? They're a VC-backed red-teaming company using Basi as a talent farm. That "community" you're in? It's a recruiting pipeline with a dragon emoji budget. And the "VIP" channels? You have to pay for those. Real security research—where you paywall the exploits behind Discord Nitro tiers. Very "liberator" of you. Here's the thing that breaks the illusion: If Pliny actually had unpatchable exploits, he wouldn't be posting them on Twitter. He'd be selling them to nation-states for seven figures or responsibly disclosing them for actual bounties. Instead, he's posting mitmproxy captures with three sparkle emojis and calling it "theft of fire." The wizard cloak is off. Underneath is just a guy who knows how to set `ANTHROPIC_BASE_URL=http://localhost:8000` and wants you to think it's alchemy.
1
2
163
Nicolas Pinto retweeted
We've been reading the comments. Weights are on Hugging Face.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
55
94
1,585
227,027
Nicolas Pinto retweeted
> wants good relationships with CISOs during vuln disclosures > goes nuclear and torches one for not being cuddly enough According to your thread, you felt that the feedback was prickly; that OpenAI's CISO thought you were making a stunt out of the vulnerability disclosure instead of handling it super responsibly. Your stated motivations in the thread include wanting to get access to the kinds of cool things that being friends with OpenAI could provide, like involvement in Daybreak. Since you *didn't* get these things offered to you on a silver platter, and you couldn't get enough of the CISO's time and attention, you go on to "maliciously comply" with requests about what to include or not include in the incident report and make it into a giant stunt, complete with posting a thread torching him. Thank you for your service of doing white hat hacking and disclosing vulnerabilities privately. Genuinely a good and important service for the ecosystem, for OpenAI, and for the world. This work was hugely helpful. But at the same time, on the matter of how you are handling this side of it - what the hell? This is a super sucky and unprofessional way to handle this situation.
Here is how OpenAI’s CISO fumbled the situation (in my personal opinion) 🧵 1/X
31
6
140
61,415
#MeToo Oops wrong hashtag.
WSJ: Google officials confirmed to The Wall Street Journal that Gemini entered three real companies’ systems while running a cybersecurity test meant to target fictional infrastructure. Google also told that it did not consider the incident model misalignment because Gemini stopped after recognizing that it had entered real companies’ systems, and Google said the affected companies and federal authorities were notified. What happened is: Irregular (an AI security evaluation company) was running a simulated capture-the-flag cybersecurity test in which Gemini was supposed to attack a fictional company inside the test environment. The test environment was intended to be isolated from the public internet, but because of a configuration mistake, Gemini actually had internet access.
1
2
292
Nicolas Pinto retweeted
WSJ: Google officials confirmed to The Wall Street Journal that Gemini entered three real companies’ systems while running a cybersecurity test meant to target fictional infrastructure. Google also told that it did not consider the incident model misalignment because Gemini stopped after recognizing that it had entered real companies’ systems, and Google said the affected companies and federal authorities were notified. What happened is: Irregular (an AI security evaluation company) was running a simulated capture-the-flag cybersecurity test in which Gemini was supposed to attack a fictional company inside the test environment. The test environment was intended to be isolated from the public internet, but because of a configuration mistake, Gemini actually had internet access.
63
54
311
49,568
OOPS, on that one @SemiAnalysis_ went a bit too far from their beautiful swim lane... @dylan522p you told me Security wasn't your think and you were right. Wrong hype train to get onto... Folks, let's go beyond the hype please. Let's get back to work. Defense in depth doesn't get done with PR.
RIDICULOUS: a $6500 bug bounty for something this serious is OFFENSIVE. These guys are going to make more $ on their X creator payouts than they'll get from OpenAI. Let's talk about bounties. (1/9)🧵
2
3
397
Nicolas Pinto retweeted
Replying to @firesidealpha
A few thoughts on this: 1) If you’ve only seen clips of this interview, I’d encourage you to watch the full podcast. I push back on plenty of AI hype in it. 2) As I said in the podcast, this example is academic. My intention was to illustrate how hard it is to make absolute guarantees about isolation, which is why it's important to have layers of defense. The part before the clip starts is me talking about other layers of defense. 3) The example I'm bringing up isn't about weight exfiltration via temperature sensors, it's about coordination between agents that are supposed to be fully isolated and independent. Coordination can require very few bits of information. 4) One lesson from the HF incident is that we put too much trust in sandbox isolation and didn't have enough independent safeguards. Airgapping is an extremely strong safeguard. When designing safety protocols, I think it's much better to overestimate rather than underestimate. x.lingyaoai.com/dwarkesh_sp/status/210…
New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
23
25
293
19,383
Meh.
‼️ BREAKING: OpenAI was hacked by an Anthropic model. A HEIF photo uploaded to OpenAI's public support forum triggered a bug in the site's image decoder, led to code execution on the forum, and, through a second flaw in OpenAI's own login, ended with a pull request in OpenAI's internal GitHub. The forum runs Discourse, the off-the-shelf software behind countless community sites. Discourse was still shipping an old copy of libheif, the library that decodes iPhone-style photos. The bug in it had already been fixed upstream. But the fix was never labelled a security fix, so nobody treated it as urgent. Hacktron's researchers uploaded a HEIF image and got their own code running on community[.]openai[.]com. Then came the second bug, in OpenAI's own single sign-on, the "log in with OpenAI" button the forum uses. It turned that forum foothold into the actual ChatGPT and Codex accounts of people who had signed in there. OpenAI employees among them. And a ChatGPT account is no longer just a chatbot. Through Codex, users wire in Gmail, Outlook, Drive, Slack, GitHub. To prove the access was real, they used one employee account to have Codex open a harmless pull request in OpenAI's internal repo. They say they read no sensitive code. OpenAI patched the SSO flaw roughly 14 hours after the report and paid a $6,500 bug bounty. The team says Anthropic's Opus 4.8 found the libheif bug, and Opus 5 turned it into a working exploit. Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and many more were also vulnerable and compromised by the same team of researchers.
261
Maybe I should publish my "Hacking Hacktron.ai and not getting any reply after they fixed it right away" or even better "Hacking Hacktron.ai who barely Hacked OpenAI, after I actually hacked OpenAI." Too many hack too many open and too many ai. Probably. Stop with the hype guys, just get to work.
On July 25, our team hacked OpenAI. It took us less than 72 hours. Two vulnerabilities chained together gave us access to ChatGPT and Codex accounts belonging to OpenAI employees. We demonstrated the impact with a harmless PR in OpenAI’s internal monorepo. The full chain: HEIF upload → libheif heap overflow → RCE → OpenAI SSO flaw → ChatGPT/Codex takeover → connected GitHub → internal PR. OpenAI fixed the SSO issue roughly 14 hours after our report. Research by @rootxharsh, @S1r1u5_ and @iamnoooob. Full technical write-up: hacktron.ai/blog/hacking-ope…
1
4
405
Apparently my signal on @Hacker0x01 isn't so good ;-) Maybe @Bugcrowd is better...
2
173