automating myself @coreauto ex Understanding the universe @xai ex hacking @aide_dev ex fb engineer ICPC WF its just code 👨🏼‍💻

Bay Area
skcd retweeted
First PSSR came to PS5 Pro. Now Quick Spectral Super Resolution (QSSR) is coming to PS5. Details on new upscaling tech, coming first to Marvel’s Wolverine and Ghost of Yōtei: play.st/4hjL9yo
1,030
1,306
14,862
1,854,606
the limiting factor is my own desire
4
1
68
5,695
good memories
2
25
4,859
astra is running a loop where its checking my battery and making sure I have enough charge on my laptop... it stops if the batter drops below 5% cause it might have builds running which would get cancelled??
5
61
8,981
I can’t stress enough how little an idea matters compared to the agency of the people executing the idea. I have had the privilege of knowing and sometimes even working with some of the most successful people (by various metrics). The difference between mediocre and excellent work and outcomes is predominantly one of agency. In practice this means: they dont wait for things to happen to them they go out and make things happen for them. They don’t wait for someone else to do something, for someone to teach them, for someone to give them the path, etc. They just go out and find a way to do it. I think the single biggest superpower these people have is the realization/belief that the world around them is completely mutable. Most everything that happens is because a person made it happen. I used to tell people to look around the room you’re sitting in. Look at everything. Every noun. It almost all exists because a person willed it into existence. Nothing is stopping you from doing the same. I see people online all the time dismissing someone else’s success because “I had that idea first” or whatever. I mean… yeah? If so then the difference is… you. So a bit of a self own whenever I hear that. Number one tip: act with agency.
228
1,279
11,859
685,041
If the same general agent can understand a car, an arm, a humanoid, etc. through their sensor/action interfaces, you start to decouple intelligence from the hardware platform itself. “LLM as puppeteer” is probably just the beginning. The bigger shift might be robots converging on an abstraction layer where a general agent can express intent without needing to understand the low-level control details of every embodiment. Almost like a hardware abstraction layer for physical intelligence.
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." web.mit.edu/phillipi/www/wri… I think it's an important change in the trajectory of robotics!
1
1
11
2,919
a more nuanced take about "can a LLM control a robot?" implies
I think “can an LLM control a robot?” is slightly the wrong question. With enough compute and enough time between actions, probably yes. What’s more interesting is what has to change in the architecture, training, and inference for these models to operate at robot speed. Robotics operates at multiple timescales. Deciding “Pick up the pan and place it on the stove.” can happen relatively slowly. Maintaining balance, correcting a slipping grasp, reacting to contact, or adjusting a joint trajectory cannot. A model that has to autoregressively generate a sequence of tokens before every physical correction is paying an enormous latency tax for capabilities you often don’t need at that layer. Making the LLM 2x or 5x faster helps, but it doesn’t really change the shape of the problem.
1
1
24
6,399
is astra inference okay? so many missing whitespaces in text, mangled text in general. it is mentioned in their model guidance docs (subagent communication section), but still interesting to see it happen with the main agent <> human channel as well
5
1
80
9,637
somedays you just need to look at graphs very very carefully and audit your agent's work and the results can be so shockingly good; a system which was failing on its head now has a lot more scaling capability
6
2
77
6,987
skcd retweeted
Baby model understands love.
7
2
67
26,900
I might have an understanding of what "software factory" really means and no its not workflows or /goal ...
14
2
110
13,802
> The pursuit of excellence does not need justification.
Mind boggling to me that I can make a thing faster and there's always people that ask "but why?" What kind of mentality is that? The pursuit of excellence does not need justification. Also, I find in so many cases, we can't know the impact of an improvement until we do it. For example, one I've talked about before: Ghostty's high IO throughput has enabled terminal program (emulator and TUI) fuzzing at a speed thats incomparably fast to prior solutions. This has resulted in upstream patches to resolve issues in popular projects like btop, tmux, and more. Speed enabled that anecdotally example that lifted the tides of adjacent communities that don't rely on Ghostty technology at all. I didn't predict this. Make things better because they can be better and let the results naturally play out.
3
1
82
12,280
we open-sourced the full Grok Build app for anyone to take a look at the code :) you can see all parts of how the agent loop runs, how the terminal rendering works, compaction, goals, subagents and more
We've open-sourced Grok Build and have reset usage limits for all users. Open sourcing Grok Build allows anyone to support making a reliable and robust harness. Check out our code, including the Git repo for the Grok Build CLI. x.ai/open-source
97
74
1,691
136,819
SpaceXAI’s Grok 4.5 scores 54 to place fourth on the Artificial Analysis Intelligence Index following only Fable 5, GPT-5.5, and Opus 4.8. It scores on par with GPT-5.5 in Codex on the Artificial Analysis Coding Agent Index in the Grok Build harness, at much lower cost Grok 4.5 improves 16 points over Grok 4.3 on the Intelligence Index, bringing SpaceXAI to the intelligence frontier behind only OpenAI and Anthropic, and outperforming all open weights models and notably Google’s Gemini models. Key standout areas of performance are agentic knowledge work and coding. Grok 4.5 in Grok Build scores 76 on the Artificial Analysis Coding Agent Index, on par with GPT-5.5 (xhigh) in Codex and just below Fable 5 (max) in Claude Code, and at a small fraction of the token usage and price. Congratulations to @SpaceXAI, @cursor_ai, and @elonmusk on the impressive release! Key Takeaways: ➤ Grok 4.5 performs very strongly on agentic tasks. Grok 4.5 ranks #4 on GDPval-AA v2 with an Elo of 1543, between Claude Opus 4.8 (1600) and GLM-5.2 (1513). It achieves the top score on 𝜏³-Banking of 33%, above 31% from GPT-5.5 (xhigh), and sits on the cost vs performance Pareto frontier across all three agentic evaluations in the Intelligence Index ➤ Grok 4.5 is one of the most cost efficient models to run for near-frontier intelligence. It costs $0.31 per task on the Artificial Analysis Intelligence Index and $2.59 per task on the Artificial Analysis Coding Agent Index within Grok Build ➤ Low cost for Grok 4.5 is driven by both low pricing and token efficiency. Grok 4.5 has a headline price over 60% lower than Claude Opus 4.8 and GPT-5.5, and used ~14k output tokens per Intelligence Index Task - over 60% lower than Opus 4.8. On the Coding Agent Index, Grok 4.5 stands out on the Pareto frontier of Coding Agent Index score vs. Total Tokens, using only 1.9M tokens for the Coding Agent Index while scoring 76 ➤ As a coding agent, Grok 4.5 in Grok Build is on par with GPT-5.5 and offers efficiency benefits: In our Artificial Intelligence Coding Agent Index that consists of DeepSWE, Terminal-Bench v2, and SWE-Atlas QnA, Grok 4.5 in Grok Build ranks third, on par with GPT-5.5 (Codex) and below Fable 5 (Claude Code). It is also very efficient in achieving this result: Grok 4.5 in Grok Build cost $2.49 per task while Fable 5 in Claude Code cost $11.80 and GPT-5.5 in Codex $5.07. This is driven by relatively low token pricing and the model using far fewer tokens than comparable models (1.9M average tokens used per task), significantly less than Fable 5 in Claude Code (7.2M) and GPT-5.5 in Codex (6.2M) Other model details: ➤ Context window of 500k tokens - a reduction from Grok 4.3’s 1M token context, but retaining configurable reasoning and vision input ➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits are discounted by 75% to $0.5 per 1M tokens, and costs still double with long (>200k token) inputs ➤ As Elon Musk has disclosed, Grok 4.5 is 3x larger than its predecessor at 1.5T parameters
108
254
1,823
1,994,680
Now that Grok 4.5 is out, here are some of my workflows which I use daily 1. "use as many subagents and tokens as you need" I end my prompts with the above to tell grok to spin up multiple subagents at the same time to divide and conquer on a hard problem. Grok is smart enough to realize when a subagent needs its own worktree and if it should inherit context from the main agent.
45
54
1,046
57,292
3. `/loop`: the same prompt, on a cadence Intervals go from 60 seconds to days (`30m`, `2h`, `1d`), recurring loops auto-expire after 7 days, and you can cancel any loop by its job ID. It's the perfect shape for **watching things**: CI logs, a canary deploy, a live production service, a flaky test you're trying to catch in the act. The magic is that each firing is a full agent turn, not a dumb cron job. The loop doesn't just fetch the logs — it *reads* them, compares against what it saw last time, and decides whether anything is worth telling you about.
3
1
90
6,909
4. `/dashboard`: control a fleet, not an agent The dispatch box at the bottom always spawns a *new* session - type a prompt, hit Enter, and you've launched another agent without leaving the screen. Do it five times in a row. Select any row to get a **peek panel**: the agent's latest output plus a live reply input, so you can answer a permission prompt with a single keystroke (`1`–`9`) or queue a follow-up without ever opening the full conversation. `Ctrl+S` dispatches *and* attaches when you do want to dive in.
2
1
82
5,052
skcd retweeted
Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency. x.ai/news/grok-4-5
1,493
3,624
26,864
69,279,491
Grok 4.5!
49
46
1,034
67,401