Building sovereign AI infrastructure. 15-node cluster · 6× RTX PRO 6000 Blackwell (576GB VRAM) · 2TB ECC RAM · PB-scale Ceph · 200GbE fabric

Mountain View, CA / London, UK
Pinned Tweet
People are hitting usage limits while dictating. Your agent's built-in voice input streams your audio to a server and back. Metered, cloud-dependent, account-locked. #NVIDIA #Parakeet v3 runs locally on any #Apple Silicon #Mac. Zero network latency. Unlimited. Free forever. Paste this into Claude Code or Codex 👇
Made with AI
3
5
374
#elonmusk is a human Drake equation , a quantum fluctuation from an #improbability drive. Shouldn’t be part of this timeline tbh
1
2
46
People are hitting usage limits while dictating. Your agent's built-in voice input streams your audio to a server and back. Metered, cloud-dependent, account-locked. #NVIDIA #Parakeet v3 runs locally on any #Apple Silicon #Mac. Zero network latency. Unlimited. Free forever. Paste this into Claude Code or Codex 👇
Made with AI
3
5
374
Set up system-wide local push-to-talk dictation on this Apple Silicon Mac using NVIDIA Parakeet TDT 0.6B v3 via parakeet-mlx. 1. brew install ffmpeg; uv tool install parakeet-mlx (install uv/brew if missing). 2. Write ~/.local/dictate/dictate.py: - Load mlx-community/parakeet-tdt-0.6b-v3 ONCE at startup and keep it warm in memory. - Global hotkey: hold Right Option to record (sounddevice, 16kHz mono), release to transcribe. - Paste the result at the cursor: save the clipboard, pbcopy the text, send Cmd+V, then restore the clipboard. - Log errors to ~/.local/dictate/dictate.log. 3. Install it as a launchd LaunchAgent so it starts at login and restarts on crash. 4. Tell me exactly which macOS permissions to grant (Microphone, Accessibility, Input Monitoring) and to which binary. 5. Test end-to-end and report the transcription latency.
51
GPUs on the operating table. RTX PRO 6000 Blackwell cards out of the ESD bags and into our own cluster. Thanks to @NVIDIAAI and the NVIDIA Inception programme for backing a small team that wants to own its compute. Benchmarks once it's racked. #NVIDIAInception #GPU
2
9
200
Every frame, sound and word of this film is code. No cameras, stock or music files. My idea. @claudeai agents built it: a Blender world, simulated basalt, narration, score. Our @nvidia GPUs rendered it in ~1 hour vs ~12 on CPU. @AnthropicAI Þingvellir, where the plates part.
1
1
9
688
normal people: I wanna be rich elon: Kardashev 0 → 1 → 2 → 3 no problem
Replying to @minchoi
Colossus 1 is 150k H100, 50k H200 and 30k GB200. Colossus 2 is 110k GB200 and 440k GB300. Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.
3
13
364
DeepSeek-V4.1-Flash running on our own 4x RTX PRO 6000 Blackwell, split across two boxes over 200G RoCE. No NVLink. 178 tok/s single-stream, 642 tok/s at 8 streams, 7-8k tok/s prefill. Engram tables stay on NVMe. Recipe and full numbers below. #DeepSeek #AI #NVIDIA
2
15
1,202
Stack: SGLang preview + SM120 FlashMLA patch + Engram-on-NVMe row store (0xSero launcher), TP4/EP4 across two nodes, DSpark block 5, custom all-reduce disabled. Native MXFP4 checkpoint, 475 GiB. NVFP4 repacks buy nothing: the experts are already 4-bit and no engine loads them.
1
8
180
How we split 4x RTX PRO 6000 (384 GB) on-prem: Normal: GLM-5.3-Flash EXL3 on 2 cards as the resident agent. Nemotron 30B, embeddings, reranker, OCR, Whisper, TTS on the other 2. Big-brain: one command, GLM-5.3 755B at 3.25 bpw takes all 4. 36 tok/s. Reverts itself. #PureTensor
1
2
2,406
GLM-5.3-Flash's NoPE MLA (rope_dim=0) breaks every stack's hardcoded 64-dim assumption, vLLM, SGLang, TRT-LLM, FlashInfer all refuse it on SM120. What worked: llama.cpp's day-old glm5next PR, own FP8→Q8_0 convert, 4× Blackwell over 200G RDMA RPC. 46 tok/s, 262k ctx, day zero
1
384
The legend himself @bcherny & his bodyguard @claudeai @ClaudeDevs code in London. #ClaudeCode #AI #DevTools
239
Visited AMAX HQ in Fremont last week. Deep technical discussion and a tour of their lab. A real glimpse of the next generation of NVIDIA supercomputers being built. #AMAX #NVIDIA #AIInfrastructure
111
Runaway agents and malicious prompt injections are a permissions-and-backup problem, not a model problem. Layered permissions bound what an agent can touch. Scoped creds, no admin keyring, no push channel to the offsite mirror, no path to the sealed seed. Layered backups recover what an agent can destroy. Snapshot lattice, immutable cold archive with object-lock, pull-based geo-mirror, verified restore drills. The recovery surface is strictly larger than the destruction surface. By design. Whatever an agent can reach to destroy is by definition recoverable from what it cannot reach. That is the threat model. There is no third thing. The Cursor/Railway incident was single-credential, single-volume, single-site. Identical outcome from a drunk intern or a rushed migration. The agent is incidental. Solved problem. Has been for thirty years. #AgenticAI #PromptInjection #AISecurity #InfoSec x.lingyaoai.com/lifeof_jer/status/2048…
142
Third RTX PRO 6000 Blackwell Max-Q arrived for the experimental Trinity cluster at PureTensor, self-imported from the US via @nvidia's Inception channel. Thanks Sean Ardura and the Inception team for the swift reseller connection.
2
2
306
Now 3 Blackwells across 2 nodes on 200G RoCE with GPU Direct RDMA. First inter-node run: Llama 3.1 405B AWQ-INT4 pipeline-parallel under vLLM on Ray. 126 layers, 42 per stage. NCCL all-reduce held 23 GB/s cross-node, 99% of wire rate.
1
101