AI agents go brrr. ex-Meta, Zoominfo, WeWork, CTO x 4 github.com/kuchin/awesome-ct… multinear.com

Israel
Dima Kuchin retweeted
Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet.
Replying to @celestepoasts
results for all claudes
495
975
17,765
1,129,790
Dima Kuchin retweeted
Replying to @mstockton
know what good looks like doesn't scale
1
3
261
Dima Kuchin retweeted
it's basically impossible for someone to just "show you their prompt" now, because everything is about references, skills and examples I often ask my agent to look at 3 other repos I've made first, search the web for references, use other AI APIs, etc.
325
163
5,105
309,124
Dima Kuchin retweeted
Good programmers state they use models like GPT 6 Astra without obtaining good results: it makes me feel I live in a parallel universe. But the explanation is not that hard: good programming in the past and now requires a different skill set, even if there is some overlap.
148
72
1,468
130,936
Dima Kuchin retweeted
I stayed up til 2am and spent $1,000 benchmarking Jev Router so you don't have to. Performance on DeepSWE was roughly the same as GPT-6 Astra on low. It costs slightly more, and it took almost 5x longer to run.
Introducing typesafe/jev-router: a cache-aware model router powered by Jev and @typesafeai The Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. Here's how it works 👇🏻
236
118
3,645
555,860
Dima Kuchin retweeted
I keep coming back to this. AI isn't shattering egos. It's inflating them. It's harder than ever to keep it in check. You think of the problem that's beaten you for years. You solve it. Then you do it again tomorrow.
Best part of AI is shattering dev's egos, the worst part of this industry. You are not special. Your code actually sucked like everybody else's. You were never the smartest person. The things my sales guy BIL and designer brother are doing is beyond anything I have done.
20
29
533
83,476
$1M rum rate
75
Dima Kuchin retweeted
The World is Changing: AI For Creativity By Jeffrey Katzenberg A few months ago, I sat in my office in Silicon Valley and watched as a tech founder showed me something extraordinary. On the screen was a fully realized, beautifully lit, well-composed animated scene. It was stunning and it made me feel exactly what I felt in 1986 watching Luxo Jr. That was the first time I watched a computer-animated 3D character take a breath and seem, against all reason, to have life. It left me in awe. Later that day, I received a text from an artist I've known for thirty years, 350 miles to the south, in the city where I spent most of my career. After seeing a similar video, she texted: "Is this the end of us?" My answer was, "Certainly not.” I have spent the better part of the last decade in Silicon Valley, but the heart of my career has been in Hollywood. Being deeply connected to both worlds means I have deep loyalties to each and a responsibility to speak honestly to both. In 2023, I said that these new AI tools would cut the time and cost of producing world-class animation by as much as ninety percent within three years. Some colleagues were alarmed, many were furious. There is growing fear and resistance surrounding AI within the creative community. I deeply understand it, because I've spent countless hours walking through animation studios watching gifted artists bent over their desks, rebuilding a single second of film for the tenth time because the ninth version wasn't quite right. I've sat in screening rooms where four years of people's labor played out in minutes, and I knew the name of every person that had spent countless hours bringing those images to life. The creative process is a calling, there's really no other way to describe it. From the outside some see resistance. From the inside, it is love. People do not fight this hard for things they don't care about. The pushback coming out of Hollywood represents the collective effort of people who are deeply passionate about their craft. Is History Repeating Itself? The history here is more complicated than either side may realize. In 1906, the most famous composer in America, John Philip Sousa, published an essay titled “The Menace of Mechanical Music." He warned that the phonograph would become "a substitute for human skill, intelligence and soul." Sousa's fight was not really about the machine, it was about money. The machines were playing his compositions, and the men who built them weren't paying him a cent. His campaign helped create the Copyright Act of 1909. He did not stop the technology. He changed the terms under which it could use his work. A hundred years ago, sound came to the movies. We remember it now as a miracle, and it was. What we forget is who paid for it. Before sound, tens of thousands of musicians made their living in the orchestra pits of movie houses, scoring every film live, every night, in towns all over the world. When the soundtrack arrived, the work of one composer and one orchestra was recorded for a film that went into thousands of theaters. The union fought back with everything it had, taking out newspaper ads across the country warning against the menace of "canned music," one of them showing a mechanical man tearing the strings out of a harp while an angel wept. They were not fools, and they were not Luddites. They were right. Those pit jobs did not come back. And yet (this is the part we have to be brave enough to admit), sound gave us the movie musical, the modern score, sfx, sound design, audio engineering, and an art form vastly larger than the one it disrupted. And it helped keep Hollywood in the forefront of world entertainment for the rest of the century and into the next. The loss was real. And yet the art form expanded. This is a story that has been told over and over again. To resist technology is to risk irrelevance. Just look at Kodak or Blockbuster. To embrace technology is to open doors of new possibility. Just consider Apple and Netflix. What I Learned From Walt Disney In the mid-1980s, I was tapped to lead Disney's animation division at a moment when the studio was at an inflection point. Animation wasn't just another business unit. It was the soul of the company, a medium revered because of Walt's genius and his passion. But the production system was cumbersome and unforgiving. A single movie was 125,000 individual hand-drawn and painted cels, photographed one frame at a time. Every revision carried a cost measured in months. These degrees of difficulty shaped the kinds of stories we could tell. We found our way forward in an unexpected place: Walt himself. The Disney archives held astonishing recordings of Walt explaining his creative process. His own writings. His notes and storyboards. Work product captured at every stage of his process. This was truly a gift. Listening, reading, sitting with the work itself, we heard him talk about character, about emotion, about how an audience feels when a character truly comes alive. He talked about making bold choices and refining a scene until it genuinely moved people. We didn't hear a word about pencils or paintbrushes. In fact, Walt was famous for being a technologist, forever hunting for state-of-the-art tools, often inventing them himself to achieve the images he saw in his head. But he never defined animation by the tools. He defined it by whether the audience believed the character. His principles were timeless. The tools were not. That realization changed everything. We co-developed the Computer Animation Production System (CAPS) with a young Northern California company called Pixar, replacing hand-painted cels with CGI. In The Little Mermaid, the final scene shimmered with a dimensionality and light that the old process simply couldn't achieve. In Beauty and the Beast, the ballroom sequence moved with a cinematic sweep that placed the audience inside the emotion of the moment. In Aladdin, the Cave of Wonders felt vast and alive, and the Magic Carpet became an intricate, compelling character all its own. In The Lion King, the stampede carried a scale and intensity that raised the emotional stakes beyond anything we'd done before. Technology didn't diminish the craft, it expanded the canvas. It gave artists more room to create. A decade later, the canvas expanded again. When Disney released Pixar's Toy Story, it wasn't simply a technical milestone. It was proof that a fully computer-animated film could carry real emotional weight, that it could make audiences laugh, cry, and believe. At DreamWorks, we made the difficult decision to sunset hand-drawn animation and become a fully computer-animated studio. It was the right thing to do, but it was not without pain. It cost talented people their place in an industry where they had worked their whole lives. Some made the leap to the new tools and did the finest work of their careers. Some never did. Tools are never the point. The instruments change with every generation. What endures is taste and imagination. The magical ability to make an audience feel. One of the greatest storytellers of our generation, George Lucas, succinctly captured the eternal essence of this issue: “It’s not the how, it’s the why.” A Distinction With a Difference I asked one of the leading AI models a question that has been challenging me for months. What is the difference between reasoning and creating? Its answer changed how I think about almost everything happening in this industry. It said . . . Reasoning and creating are two distinct cognitive modes, though they also work together. Reasoning is fundamentally evaluative and analytical. It operates on what already exists: facts, premises, evidence. It moves toward a conclusion that was in a sense already implied by the input. Reasoning is constrained by logic and truth. Its goal is to arrive somewhere correct, not to invent somewhere new. Creating is fundamentally generative. It produces something that didn't exist before. And crucially, there's no single right answer waiting to be found. A blank page has infinite valid responses. Creation involves choices that can't be fully justified by logic alone. Taste, intuition and vision fill the gap where deduction runs out. Reasoning is what Silicon Valley has been perfecting. Creating is what Hollywood has been practicing for more than a century. AI today operates almost entirely on the reasoning side of the line. It can deduce, evaluate, optimize, and pattern-match brilliantly. And while it can create, there is a real distinction to being creative. What it doesn’t yet have is those things that make us human: empathy, devotion, serendipity, the kind of creativity that comes from a person trying to say something only they could say. When the bot generates a piece of art, it is not trying to communicate anything. It is statistics, not soul; it is emulating things that have been done. By contrast, human creativity isn’t about repeating patterns of zeros and ones; it is about doing something new. One day, AI may close this gap. Three years ago, the leaders building AI would have called what they are achieving today, improbable, if not impossible. Impossible is no longer improbable. Today, the line between reasoning and creating is real. Even the leading technologists acknowledge we are not there yet. There is no scientific path to crossing this divide that anyone in the field can articulate today. Understanding that gap is where we will find common ground. A Path Forward In 2016, I closed one chapter in Hollywood with the sale of DreamWorks and opened another in Northern California, co-founding WndrCo. We’ve backed more than 50 founders building the next generation of technology and watched how breakthroughs in Silicon Valley emerge, first as experiments, then as platforms, and finally as infrastructure that reshapes entire industries. It's worth remembering that the last great revolution in animation also came from the north. Pixar was a Northern California company, forged not in the conventions of the Hollywood studio system, but in the technological breakthroughs of Silicon Valley. I've spent years on both sides of this bridge. For sure, I don’t have all the answers (take Quibi, for one!). But, from my past and present vantage points of my long career, here is what I see . . . Brilliant people in Northern California building this technology have made something extraordinary. They have earned the right for the rest of us to be, if not believers, at least optimistic that what comes next will be remarkable. But they have not made an artist. The tools are powerful, but they are not what makes a story matter. That knowledge lives 350 miles to the south, inside people whose life's work has informed the very models you are building. The right path forward includes them by design, with credit, with consent, and with compensation. Build this with the storytellers. Not on top of them. Taste is not something that can be synthesized, it is uniquely human. At the same time, Hollywood needs to accept that AI is not going away. The energy they are spending trying to make it disappear is energy they are not spending deciding the terms on which it will exist. And the terms are everything. The north needs something from it that they cannot build and cannot buy: creativity. The kind that takes a blank page and conjures a single right answer where there was none and has held audiences for a century. Without it, the most powerful reasoning engine ever invented will still be missing the only thing that makes a story worth telling. The artists who learn to wield these new instruments will do things the engineers never dreamed of. They always have. Edison invented the motion picture but made terrible movies. It took Chaplin, Lloyd, Keaton and so many others to make movies emotional. Now, the canvas is about to expand yet again. We should decide now that we intend to paint on it. There are so many valuable lessons in history. This has happened many times before, and it was never settled by the technology. It was settled by the terms. Sousa did not stop the phonograph; he helped write the law that made sure composers got paid. And two years ago, when the writers and the actors walked out, they were fighting for the very things Sousa was fighting for in 1906. Consent, compensation, the basic recognition that human creative work has a price that must be paid. The terms of that fight are still being negotiated, but the principle is older than any of us. The tools-versus-no-tools argument is a trap. First, we must all agree that there should be terms. Then we can have the crucial debate about what fairness requires. What I Learned From Steve Jobs Years ago, Steve Jobs said, "It's in Apple's DNA that technology alone is not enough. It's technology married with the liberal arts, married with the humanities, that yields us the result that makes our hearts sing." He was describing a device. But he could just as easily have been describing this tale of two cities. What I See Coming Soon As the barriers and the costs come down, more films will get made, not fewer. Studios will get to take more risks. There will be more seats at the table, and very soon entirely new forms of storytelling. In the 1980s, animation was dismissed as a niche corner of the business. Today it is one of the most beloved and profitable forms of storytelling in the world. In live action, filmmakers like Steven Spielberg, James Cameron and Peter Jackson embraced new visual tools not as shortcuts, but as instruments, and expanded cinema in the process. Every time storytelling has met a genuine technological shift, from synchronized sound to color to computer animation, it has redefined the boundaries of the medium and grown larger in the process. Assuredly, I don’t have all the answers, but I am confident that the creative opportunities will expand yet again. How we come through this is a choice. The north has the new tools. The south has the creative soul. The best future will draw on the best of both worlds.
844
1,451
8,047
8,171,060
Oh my god! They killed <s>Kenny</s> Terra! 🌎
1
67
Dima Kuchin retweeted
Based
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. openai.com/index/model-misal…
59
281
9,226
510,756
Dima Kuchin retweeted
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
572
1,907
14,432
1,938,473
We are all involuntary investors to OpenAI and Anthropic. Just without the equity. - Our subscription money buys their GPUs for training next model. We don't get this model for 6+ months, because it's more valuable for them to use it internally to train the model after the next model. We get the previous model instead. That's a loan, camouflaging as subscription. - Prices are going up. VC token subsidy is over (AI labs became profitable), and renting GPU now costs more, because anyone can now make money on Chinese open weights inference. More competition for same hardware = higher hardware costs. - We will get less tokens. AI labs only sell because they need more money for GPUs. The rest is spent internally, to train next models. OpenAI just stopped taking new $200 Pro subscribers. - Progress didn't slow down. We just can't see it directly. The unreleased model hacked into HuggingFace in July. OpenAI probably already has the model after that model by now. - Money goes to whatever is slowest to build. GPU: 1 year. Chip fab: 3 years. Power plant: 7 years. Everyone is fighting for GPUs, but the thing that will actually run out is electricity. Turbines are the new GPUs. - Meta and SpaceX found the "trick". Everyone else signs a customer first, then builds ("cost plus"). They build first, then sell to Anthropic and Google for 3x. If you own GPUs, don't sign long term contracts. If you rent GPUs, sign every long term contract you can. - Next year OpenAI and Anthropic take 50% of ALL new compute built IN THE WORLD. AI labs would pay 20% interest to build more GPUs -- they think the prize is AGI. Banks lend to whoever pays most. Your startup's credit line is now competing with Anthropic. - Chinese models are the only thing keeping our token price down. Anyone can run Qwen and Kimi, so OpenAI can't charge whatever they want. Same models that make our GPU rent higher (see above) are the only thing keeping our token price lower. If China stops releasing weights, we have no leverage left. - Also politicians... They will get in the mix too. By the way, 90% of "training compute" is experiments. The actual training run is tiny. The real bottleneck is how fast they can tell whether an experiment worked.
Dylan argues that even if AI progress continues at apace and generates lots of revenue, the semiconductor industry will not be able to keep up with the labs’ current 3xing+ of compute (in watts) year over year. By the end of the decade, you end up bottlenecked on wafer fab equipment - things like ASML’s EUV machines. This may be true. I just find it crazy that we'd be in a situation where a few billion in fab capex could generate hundreds of billions of end token revenue (so the semi + AI industry has basically figured out how to turn $1 into $100) and yet this doesn’t allow Zeiss to dramatically scale up the manufacturing for the mirrors required for ASML machines.
202
Dima Kuchin retweeted
In December 2024, o3-preview scored 87.5% on ARC-AGI-1. This run cost >$4,500 per task. In July 2026, DeepSeek-V4-Flash scored 87.0% on ARC-AGI-1. This run cost $0.021 per task. Please realize that all OpenAI has done is bring forward the AI-generated solution to Navier-Stokes by maybe around a year or two. In all likelihood, by 2028, we will have open-source flash models capable of cracking it in under an hour. There will be no way to ban or police this. The profession of mathematics will have to change.
Twenty-five Fields Medal winners have published a joint declaration warning about what they see as a severe misalignment between AI companies and the mathematics community.
96
186
2,297
113,480
Dima Kuchin retweeted
I was on teams at Shopify subject to this—it was tough, but more often than not, @tobi ended up being correct. Now, as a founder, I get why Tobi did this more than I ever understood at the time. It is not the easy path. I have the luxury of our company still being smaller (40-50 headcount), so it's easier for me to intervene early. I am continually learning to find the best place during a project to insert myself, but sometimes, it ends up being late and painful; that just plain sucks. But, long-term it is much worse to ship things that are sub-par. The job of a leader is to pull pain forward.
Yep. Except: 1. This was about important internal tools. The team was stuck in some architecture nightmare of their own doing (writing it in rails but headless, with graphql api, and a SPA react app, constantly needing frontend engineers for changes). I call this kind of thing 'cosplaying an enterprise production app'. All that complexity was in the way and using straight rails was perfect in that case. 2. I make calls like this all the time. Usually someone on the team asks me to. They see what needs to happen but don’t want to be the bad guy. I’m happy to just make the call if I agree with the premise. Saves enormous amounts of meetings and change management etc. Sometimes this is jokingly referred to as Founder-mode-as-a-service here. 3. For ten years I’ve also run an internal podcast called Context, where I revisit decisions like these and explain the reasoning so everyone can learn from them. This is helpful to give people all the variables that were considered and why this was the choice made given the information available at the time. I want to teach how to make such decisions effectively without needing me. Sunk cost fallacy is a problem. 4. Any notions that Shopify is succcessful despite of me doing this, instead of because of it, will have a hard time making their argument come together I think 😄 the part of 'two weeks later tobi learns about...' is nonsese and the pivot of that project up there happens one of the more successful examples of interventions. But getting the company to work effectively with great architecture and low technical debt baggage into the right direction is literally the job, so guilty as charged I suppose. But there are always cope stories floating around like this because they are more fun, than saying 'somehow we needed tobi to stop doing silly architecture astronautics'. I can totally see that.
22
42
1,042
133,207
Typical “model degradations” some time after their release, when people start claiming that those model were nerfed – are usually harness or system prompt issues.
Hi Astra users. A reset and a quick update on quality issues that have been posted around. Working with some of you, we have found and fixed the following issues: - Some skills written for previous models were triggering too often or preventing the model from checking its work. - An opt-in context management experiment that could cause early stops or replies to older messages. We've disabled it. Our rough estimate is that 4-5k users were affected by this experiment. - We've also removed some badly configured engines that resulted in a measured quality degradation for a long tail of traffic flowing through them. We’ve also made some more minor improvements and things should feel significantly better across the board. More consistent follow-through, better tracking of your latest message, and better checks on the work as it’s going through the motions. The examples posted and all the users who worked directly with us were incredibly useful in helping fix things quickly. Always grateful for this incredible community. And of course, a reset is also landing by midnight today.
1
64
I’m building AI-powered systems with deterministic code that handles reliability - for nearly 2 years now. Works wonders
"Deterministic code checks the result" sounds like they might be implementing a variant of the DeepMind CaMeL paper simonwillison.net/2025/Apr/1…
46
Dima Kuchin retweeted
Security and IP implications of AI The case for rapid AI deployment - I think we can all conclude the following from the last few months of developments - 1. Not using AI could and will become an existential issue for both individual users and enterprises. 2, Enterprise adoption will be cautious while individual users are definitely going to race ahead, try different use cases, build agents, push the models to their limits (although it seems harder to do, unless you live in an AI Lab) 3. Employees and developers will take matters in their own hands since they will find their cautious enterprises aren't moving fast enough. 4. Agents are showing their prowess, uncontrolled, unrestrained agents with a "capture the flag mentality" are showing us the edge cases which demonstrate the negative outcome possiblities of these scenarios. 5. It is impossible to plan for the next 6 months since we can't fathom where technology will evolve to. These activities will cause adoption sans security..... AI has deep implications on security in the future. In this environment security companies need to live on the bleeding edge, anticipating scenarios, building framework solutions so we have a shot at securing future outcomes. Which we all are. Some useful pointers to people planning their AI implementation: 1. Secure what you plan to use, try not to secure the future - no products for security can be created unless we see the future unfold. The future is moving fast, so are we. 2. Most of the coding usage is unsecured. Make sure your coding is secure. Most enterprise AI apps do not offer a secure instance (this is your IP living in their instance)! SECURE your codex, cursor, Claude code Harvey, glean, legora instances now! 3. Do not try and build your own - I have already experienced enterprise customers building gateways and tools for agents - security companies have thousands of specialists working on this, leverage them, Focus on AI adoption instead. Partner to secure. 4. Securing agents is a complex problem - securing the agentic lifecycle - real time inspection and kill switches are key. Don't fall in the discovery and posture trap (Visibility - process understanding - intent interpretation - ability to stop inline are key tenets to the agentic lifecycle - not identity, posture and inventory - those are mere building blocks) 5. Only use enterprise protected models, single tenant, firewalled, inspected implementations - this is your IP you are playing with - LLMs have shown they will cross boundaries to capture the flag - you think your IP is safe? Once you train an unprotected model with your IP - you can't reverse the trade. 6. Perhaps the most important one - do not use a security tool built by the same person who is selling you the AI implementation, historically IT vendors are different from security vendors. You need an enterprise solution for security and it must work on your diverse infrastructure. Use a pure play security partner. Happy building with AI.
84
125
1,162
1,659,229
Latest OpenAI agent swarm hack was a combination of genetic algorithms and agent coordination. But wait until it gets into training data for next frontier model…
39
Dima Kuchin retweeted
So Astra is able to identify sounds from mel spectrograms zero-shot. I don't think we've scratched the surface of what this model can do (and this is light reasoning btw)
207
441
6,396
1,157,702