People are hitting usage limits while dictating.
Your agent's built-in voice input streams your audio to a server and back. Metered, cloud-dependent, account-locked.
#NVIDIA#Parakeet v3 runs locally on any #Apple Silicon #Mac. Zero network latency. Unlimited. Free forever.
Paste this into Claude Code or Codex 👇
People are hitting usage limits while dictating.
Your agent's built-in voice input streams your audio to a server and back. Metered, cloud-dependent, account-locked.
#NVIDIA#Parakeet v3 runs locally on any #Apple Silicon #Mac. Zero network latency. Unlimited. Free forever.
Paste this into Claude Code or Codex 👇
Set up system-wide local push-to-talk dictation on this Apple Silicon Mac using NVIDIA Parakeet TDT 0.6B v3 via parakeet-mlx.
1. brew install ffmpeg; uv tool install parakeet-mlx (install uv/brew if missing).
2. Write ~/.local/dictate/dictate.py:
- Load mlx-community/parakeet-tdt-0.6b-v3 ONCE at startup and keep it warm in memory.
- Global hotkey: hold Right Option to record (sounddevice, 16kHz mono), release to transcribe.
- Paste the result at the cursor: save the clipboard, pbcopy the text, send Cmd+V, then restore the clipboard.
- Log errors to ~/.local/dictate/dictate.log.
3. Install it as a launchd LaunchAgent so it starts at login and restarts on crash.
4. Tell me exactly which macOS permissions to grant (Microphone, Accessibility, Input Monitoring) and to which binary.
5. Test end-to-end and report the transcription latency.
GPUs on the operating table. RTX PRO 6000 Blackwell cards out of the ESD bags and into our own cluster. Thanks to @NVIDIAAI and the NVIDIA Inception programme for backing a small team that wants to own its compute. Benchmarks once it's racked. #NVIDIAInception#GPU
Every frame, sound and word of this film is code. No cameras, stock or music files.
My idea. @claudeai agents built it: a Blender world, simulated basalt, narration, score. Our @nvidia GPUs rendered it in ~1 hour vs ~12 on CPU. @AnthropicAI
Þingvellir, where the plates part.
Colossus 1 is 150k H100, 50k H200 and 30k GB200.
Colossus 2 is 110k GB200 and 440k GB300.
Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.
DeepSeek-V4.1-Flash running on our own 4x RTX PRO 6000 Blackwell, split across two boxes over 200G RoCE. No NVLink. 178 tok/s single-stream, 642 tok/s at 8 streams, 7-8k tok/s prefill. Engram tables stay on NVMe. Recipe and full numbers below. #DeepSeek#AI#NVIDIA
Credit where due: the SM120 FlashMLA patch and the Engram-on-NVMe row store come from @0xSero's launcher, github.com/0xSero/deepseek-v…. We only added the two-node split. Go look at the repo, the measurements there are meticulous.
How we split 4x RTX PRO 6000 (384 GB) on-prem:
Normal: GLM-5.3-Flash EXL3 on 2 cards as the resident agent. Nemotron 30B, embeddings, reranker, OCR, Whisper, TTS on the other 2.
Big-brain: one command, GLM-5.3 755B at 3.25 bpw takes all 4. 36 tok/s. Reverts itself.
#PureTensor
GLM-5.3-Flash's NoPE MLA (rope_dim=0) breaks every stack's hardcoded 64-dim assumption, vLLM, SGLang, TRT-LLM, FlashInfer all refuse it on SM120. What worked: llama.cpp's day-old glm5next PR, own FP8→Q8_0 convert, 4× Blackwell over 200G RDMA RPC. 46 tok/s, 262k ctx, day zero
We took UC Berkeley's FreeToken (arXiv:2608.16157) for a spin four days after release: GLM-5.2, 753B params, on a single 96GB RTX PRO 6000 Blackwell. Threadripper PRO 9975WX, 503GB DDR5. Their number: 14.9 tok/s. Ours: 16.9. @Andy_ShuoYang#AI#LLMpuretensor.ai/blog/four-days…
Visited AMAX HQ in Fremont last week. Deep technical discussion and a tour of their lab. A real glimpse of the next generation of NVIDIA supercomputers being built.
#AMAX#NVIDIA#AIInfrastructure
Runaway agents and malicious prompt injections are a permissions-and-backup problem, not a model problem.
Layered permissions bound what an agent can touch. Scoped creds, no admin keyring, no push channel to the offsite mirror, no path to the sealed seed.
Layered backups recover what an agent can destroy. Snapshot lattice, immutable cold archive with object-lock, pull-based geo-mirror, verified restore drills.
The recovery surface is strictly larger than the destruction surface. By design. Whatever an agent can reach to destroy is by definition recoverable from what it cannot reach.
That is the threat model. There is no third thing.
The Cursor/Railway incident was single-credential, single-volume, single-site. Identical outcome from a drunk intern or a rushed migration. The agent is incidental.
Solved problem. Has been for thirty years.
#AgenticAI#PromptInjection#AISecurity#InfoSecx.lingyaoai.com/lifeof_jer/status/2048…
Third RTX PRO 6000 Blackwell Max-Q arrived for the experimental Trinity cluster at PureTensor, self-imported from the US via @nvidia's Inception channel. Thanks Sean Ardura and the Inception team for the swift reseller connection.
Now 3 Blackwells across 2 nodes on 200G RoCE with GPU Direct RDMA. First inter-node run: Llama 3.1 405B AWQ-INT4 pipeline-parallel under vLLM on Ray. 126 layers, 42 per stage. NCCL all-reduce held 23 GB/s cross-node, 99% of wire rate.