Developer, Doodler, Ai Engineer, Manga Writer ✍️ Decentralize Love & Knowledge!

On-Chain
Wallet feature β€œcoming soon” in codex mobile Customize menu
1
104
Babe wake up, he has spoken
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1
133
Want to learn about kernels, and the breakdown of local LLM utilization on your system. This is easily the best write up I’ve seen recently. Refreshing to see a human write up.. lately just slop, with a side of slop Keep a human in the loop
You do not pick a model and a GPU and call it done. You pick a file encoding, and a kernel path. The GPU follows those. Start with the loop. Text becomes tokens. Tokens move through a Transformer. Attention decides which earlier tokens matter. The runtime keeps a KV cache so the model does not recompute the whole conversation every time. Then it picks the next token and does it again. The model is not writing a whole answer in one shot. It is generating one token at a time. That loop has two phases, and they are not the same job. Prefill reads the prompt and builds the first KV cache. It is compute-heavy. That is the pause before the first word. Decode writes the answer one token at a time. It keeps rereading weights and the cache. It is bandwidth-heavy. So an RTX PRO 6000 with 1.8TB/s beats the DGX Spark with 273GB/s. That is the typing effect you feel. Long prompts punish prefill. Long answers punish decode. Long chats punish both, because the working memory grows. The inference engine is not the model. It is the traffic cop, the memory manager, the kernel dispatcher, the scheduler, the cache accountant, and the API surface. It loads the weights. It tokenizes the input. It runs the forward pass. It samples the next token. It keeps the KV cache. It streams the result. Serious engines also pick kernels. A kernel is not "the model." A kernel is a specific tensor program: shapes, layouts, datatypes, and what the silicon is actually allowed to multiply. Same math on paper. Different kernel. The file format matters. It decides what can load, what can quantize, and how fast it runs. Quantization is not one switch. Storing weights in 4-bit is not the same as doing 4-bit math. Weight quantization shrinks the model. The live context is a different thing. A Q4 sticker is not universal. The right format is the one your engine has optimized kernels for. Assuming every quantization label is portable is how people buy a 5090, download "NVFP4," and still miss out on performance. Here is the worked example. I put Qwen 3.8 27B on an RTX 5090. Two downloads. Same model. Same GPU. Both folders said NVFP4. One of them does 4-bit math for real. Weights in 4-bit. Activations in 4-bit. The matrix unit can multiply them as 4-bit times 4-bit. The other stores 4-bit weights. Then unpacks them in the kernel. Then does 16-bit math. Different kernel path. The 5090 did not choose that. The checkpoint did. Especially whether the activations are 4-bit too. If they aren't, there is no legal 4-bit times 4-bit multiply to run. The engine falls back. Quietly. The file still loads. That is the difference between an encoding you can load, an operation a backend implements, and arithmetic the hardware actually executes. They not interchangeable. This is also why prefill and decode do not get the same gift. Prefill has enough token rows to keep the matrix units busy. Native 4-bit math can matter there. Decode is still walking the weights and the cache, one token at a time. If the live state never went 4-bit, you do not get a 4-bit win on that part. The bottleneck moves. Do not benchmark "the model." Benchmark the stack you will actually run. Separate prefill from decode. One more thing people get wrong with the word Blackwell. It is a marketing name on four chips that cannot run each other's kernels. A data-center B200 is not a 5090. A 5090 is not a Spark. A Spark is not a Thor. Instruction support is necessary. It is not sufficient. A checkmark on the spec sheet is not the kernel that ran. In the inference engines article under my profile, I asked a question I still want on the wall: What quantization format has optimized kernels on my target engine? This Qwen run is that question in a box. Same GPU. Same NVFP4 label. Different kernels. The engine followed. The checkpoint decided which kernel it was allowed to launch. The file loaded but that is NOT the same as the kernel you wanted for the GPU you bought.
1
96
absence of output is not absence of action
116
If you do not actively feel like you are finding all the ways NOT to make a lightbulb, then are you even pursuing incandescence?
1
2
102
Testing swarm intelligence x system 1 thinking x recursive auto-loop targeted at specific systems with: Luna, Hermes and Minimax feels like infinite experimentation Now honing in effectiveness of the swarm’s neural network is a concept that is even more fascinating.
1
90
System 1, meet system 2
Give your app real-time decision-making with Decisions API, powered by GPT-6 Luna. Define questions and possible answers to classify content, route requests, or choose an agent’s next action. Available in limited preview.
3
5
216
Ai should not be considered the industrial revolution, that is dystopian Ai is not about turning people into cogs, it's the dawn of a new renaissance
1
5
93
Vibe coded Gameboy was actually on my to do list when i catch up on research.. wonderful @sama @thsottiaux (also.. quit pretending you do not know how the button works LOL) @OpenAIDevs
1
7
90
I CAN LET SOMEONE SIGN IN AND USE THEIR OWN TOKENS ON MY TOOLS? THIS IS INSANE.. AND AN OPEN AI MARKETPLACE. your product can be BOUGHT. Plugin grassroots companies gonna EAT
1
4
96
Romaine's demo are always fun.. what an intro @OpenAIDevs Dev day!
1
2
85
damn, having to manually type?! heaven forbid πŸ˜‚
16
New tools for @OpenAIDevs Codex Harness: fully open source. desktop app wen? Codex in Cloud: (always-on, like DOTS) Codex β€œSecurity Cloud”: continuously scans for vulnerabilities, opens PRa Agents API: including Computer Use OpenAi Private Intelligence: ZDR & Private Inference
2
96
vel β‹ˆ 🀌🧦 retweeted
Introducing dots, powered by GPT-6 Astra. Remarkably capable, always-on agents built to handle everything.
2,216
3,302
37,748
14,213,553
people thinking @OpenAIDevs Dots are a clone of @bot are radically shortsighted Do you not remember @steipete , and how quickly he was scooped up for the small opensource project you may have never heard of, @openclaw? its okay to have opinions, but why be aggressively wrong?
1
64
vel β‹ˆ 🀌🧦 retweeted
Announcing dots. Dots work 24/7 for you, learn from your feedback, have their own computer, browser and can be connected to over 4k apps in our ecosystem. They’re powered by Astra our best model yet. Included in your Pro plan, without drawing down on any of your usage. You can even call a dot while it’s working. We’re getting you started with your primary dot today and soon you’ll be able to create entire teams of them.
1,438
800
12,543
2,947,768
DOTS - Your new 24/7 personalized agent. it is built to do the things you do not want to. @OpenAIDevs have incorporated the big @steipete value add as the flag ship opening to Dev Day! To put concept simply, DOTS: Dont Over Think Shit
83
Running fanned out auto-research has helped test various slices of implementation options in parallel. Implementers: -Sonnet 5.5 -Luna xHigh -Minimax m3 -Deepseek-v4.1-Flash -Local Qwopus (@KyleHessling1 and team's Qwen 3.8-27b-flash-gguf) Reviewers: -Opus5.5 -Astra
1
1
185