Long horizon tasks have definitely gotten better. Seems to be a combination of model training and harness engineering problem to make it work. The hardest thing though is verifiability and long horizon tasks by nature are difficult to verify.
"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (link in comment)
17
Vignesh Iyer retweeted
dot-com bubble vs. a possible AI bubble. From the famous "Dean of Valuation", Professor Aswath Damodaran, of NYU Stern School of Business, “And that’s the real big difference between the dot-com boom and bust and the AI boom. We don’t know whether there’ll be a bust. History suggests there will be a bust. The dot-com boom and bust had no huge capital expenditure in that cycle. In fact, there was very little traditional CapEx, or even R&D, driving it. People started apps. They basically started going on it. This has been the biggest infrastructure run-up I think I’ve ever seen in business. You can go back and compare it to the automobile business 100 years ago. The amount of money that’s being put into AI CapEx is immense, which means that when the correction comes, the pain will be more intense. And herein lies the second problem. The dot-com boom and bust was almost entirely equity-funded. You think, so what? Well, when the bust came, those shareholders lost 60%, 70%, 80%, or 90% of their money. You felt sorry for them, but the loss was restricted to the shareholders. The problem with the AI CapEx boom is that not only is it immense, but a big chunk of it is funded with debt, and the debt is coming from private capital rather than banks. There’s a very real chance that if there’s a correction and companies start having problems, that problem is going to show up as distress and default, and that really doesn’t stay restricted. It spills over into the rest of society. I’m not saying it’s going to be 2008, but 2008 is an example of what happens when lenders overreach, when they lend money at too low a rate, and the correction comes. The pain spills over. So that is my concern with this big market illusion: the potential societal cost of having to deal with debt coming due that you’re unable to pay. It’s much more painful than your share price dropping 90% and you feeling the pain." ---- From "Excess Returns" YouTube channel, (link in comment)
22
48
236
29,946
the fact that a large network of numbers can understand something at the highest level of abstraction (language) and take it to the lowest level implementation details (code) is simply mind blowing
11
the hardest shift I have been facing is letting go of my opinions about how a system should look as a programmer and let codex/claude do its thing
16
Each by themselves seem so trivial. But putting together in a sequence makes all the difference. Loving this book. It is so information dense.
16
These mods are so addictive. Reminds me of neovim days. Hopefully we get more modifiability. Had Claude create these custom loading animations.. github.com/vgnshiyer/nowload…
You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
52
Engineers often think about which new feature does my customer need. Rarely does one think about which feature they should probably remove.
1
35
Products from founders with an engineering mindset are often Feature Factories
23
its better to rejoice in all the amazing things you can now create than mourn the flow state of writing code by hand that you have lost
23
we only automate work that we hate doing.. and that was useless work to begin with.. we would never automate what we love doing. that is the only useful work.. work that we love doing.
31
Claude is notorious to draw a sword to kill a fly
27
I brought Lumbergh to Claude Code. He keeps asking me about the TPS reports.
4
5,125
new version
21
Wasn't happy with the previous version..
I brought Lumbergh to Claude Code. He keeps asking me about the TPS reports.
1
48
Vignesh Iyer retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1,378
5,694
49,470
6,099,958
it’s pleasantly disturbing when you can’t sleep because your mind cannot stop giving you ideas
23
it's a claude code mod. bill pops up above the prompt while claude works, and sometimes the spinner switches to "preparing TPS reports…". purely cosmetic, he never touches your work. /plugin marketplace add vgnshiyer/tps-report /plugin install tps-report@vgnshiyer
130
Claude Code just brought the Neovim community back from the dead
You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
64
TIL Don't use Ultracode mode in claude code for every task. Preserve it only for complex tasks. Otherwise you will be wasting a lot of time.
1
67