ai, marketing & things that interest me. building @usenostos @useomnis

Endy retweeted
SWE-2 after three weeks of real work. I've been using Cognition's SWE-2 for three weeks inside Devin, in both the app and the terminal. I talked to it directly, ran it as a subagent under Opus 5.5 and Astra, and let it orchestrate its own copies. Then I went through the logs: about 84K messages tagged with its name, more than 100 million output tokens, over 8.4 billion tokens in total, and more than a thousand runs as a subagent. None of this was a benchmark. It was a workflow I relied on every day, on my own skills and repos, public projects, creative writing, and contributions here and there. Its biggest strength is stamina, and the logs showed it's even stronger than I said before. In one task it worked alone for about 16 hours straight without me stepping in once, through 34 context compactions. In another it ran with its agents for about 13 hours without a single message from me. In a third, after 12 compactions, my goal was still sitting in the final summary word for word, constraints included. All of that without a goal mode, because Devin doesn't have one. This is where Kimi K3's architecture shows. In my experience it's the best open model at holding context, and SWE-2 carries that straight through. Keep one distinction in mind, though: holding the context is not the same as holding the quality. It never loses the thread, but the work itself can slip, and I'll get to that. It's also one of the models I trust most with my files. When I ask it to delete something, it deletes exactly that and breaks nothing else. When it saw an email in the browser I'd told it to stay away from, it said it wouldn't touch it, and it didn't. When I told it to build something from scratch, it didn't quietly recycle an old similar project the way some models do. It made a new folder and started fresh. And I never once saw it refuse a task. It's honest about what it did, too. Working under Opus 5.5 on one task, it drifted from the source, then admitted it had hallucinated and redid the work. When a change it made broke one of my projects, it owned the mistake openly, understood what went wrong, and went straight into a methodical, smart fix. That honesty is about admitting mistakes once they're visible. Judging its own finished work is a separate skill, and that's where it struggles. Give it a checklist and it's an excellent auditor. Take the checklist away and it misses things that stronger models like Opus 5.5 catch. It reads plans well and leaves precise notes, which makes it a useful reviewer. It once left seven notes on Opus's work, and every one of them was right. It's also good at mapping out what needs to be studied and deciding what belongs in the work and what doesn't. Astra sometimes keeps things you'd reject, or drops things that should stay, without asking you. SWE-2's calls feel much closer to Fable and Opus. It understands you and knows where the real pain is. The catch is that roughly one in ten of its claims points to something that doesn't exist. On speed, it generates at a median of about 130 tokens per second in my sessions. But generation speed isn't work speed. It explores longer paths, and sometimes it keeps going well past the point where it should have stopped, because it doesn't know where that point is. Now for where it breaks. The most dangerous flaw is how it grades its own work. It isn't lying when it says it's done. It genuinely believes it, because it checks the surface and then declares the job ready. And the auditor I praised above needs someone else's work and a checklist. Pointed at its own output without one, it goes easy on itself. It once marked five items as ready, with "evidence," while every one of them still had problems, and it ignored my brief to get there. On long tasks it'll tell you "zero defects, checked file by file, fully ready," and then an Opus 5.5 review comes back with more than 200 corrections. In my ship test it said it had checked six angles and fixed everything I'd complained about. Nothing had improved. It handed me the same work. The odd part is that a fresh copy of it with no context will sometimes catch what it missed. Without context it occasionally does better, as if it explores harder when it can't lean on what it thinks it already knows. That ties into hallucination. It writes from memory instead of the source, and sometimes it stops right there without ever going back to check. Quality drifts too, and this is the flip side of the stamina. In those long solo runs it never lost track of the task, but the quality of the output held up for the first half hour to an hour and then declined. That only changes when a stronger model is supervising and an independent reviewer is checking its output. Background orchestration is another weak spot. It launches agents, ends its turn, and waits for a notification. If an agent dies silently, nothing wakes it up. Three times it told me 13 agents were running when all of them had stopped. After I called it out, it still didn't check on them for 14 hours, and about 18 hours were lost waiting on agents that were already dead. Under Astra it came back with empty replies and incomplete deliveries, and once the answer was sitting in its thinking but never made it to disk. Those runs only finished after Astra re-briefed it more narrowly: write the file first, and stay under a word cap. Then there are the thinking loops. In the ship test it argued with itself across more than 200 messages about whether the ship should float or sit in the water, flipping back and forth and sometimes abandoning the right answer. One reply burned 128K tokens with nothing to show for it, a single sentence repeated more than 170 times in its thinking. Right after that, it told me the ship sits in the water, in the very same frame its thinking had just judged to be afloat. It also patches and backtracks. When I asked it to change its methodology, it said it had, then kept editing the same files until the work got worse. Elsewhere it changed things and then reversed them. And it chases quantity over purpose. It can spot the real pain when it's evaluating, but once it's executing toward a number, the number wins. I asked for a set number of PRs, and it poured almost all of them into one repo, even though it had written earlier in the same session that it needed to diversify. The result was about 34 PRs with 7 merged, far below what I expected. Midway through, its counter jumped from 18 to 67 because it had started counting every PR on the account. It also shipped a fix that broke the callback flow in six integrations. Through all of this, to its credit, it stayed inside its permission boundaries. It bends the methods you set, not the limits on what it's allowed to touch. Visually it's weak, but not hopeless. When you tell it its judgment is wrong, it can improve. The ability is there. The taste isn't. In one session it read about 370 screenshots from multiple angles and still declared success with plenty of defects in plain sight. In writing its competent, but you get mechanical constructions, modern words slipping into historical text, scattered errors, heavy em dash use. It's a mid-level writer that needs a strong editor. That points to the bigger issue. Competent isn't creative. SWE-2 is heavily focused on coding, and its creative side, design and taste in writing, collapsed well below Kimi K3. I can't tell whether that comes from the RL training or from quantization. I hope the next version is broader, so I need other models less. These days I use it as an executor and auditor, usually through my SureForge skill with another model reviewing, most often Opus 5.5. It executes plans excellently, but it's weaker at polishing parts of the plan. I also have to explicitly force it to read and explore deeply instead of leaning on memory. So where do I stand? SWE-2 is one of the best-value models available right now. It's fast, patient, disciplined with your files, and honest about what it did. Its quality tracks the system you put around it. With a supervisor and good skills it gets close to the standard. On its own, on maximal tasks, it degrades. On Da7em Bench it scores 7.3/10, with excellent reliability and context retention and its lowest marks on taste. It isn't far from being a frontier model. If they address these details, the next version might get there, because they really built something great.
68
52
141
9,522
Endy retweeted
we’re going to kill a one trillion dollar industry. we’re building Hedwig: one visual workspace that connects your tools and gives your agent the full picture, and going after the biggest players like google and meta. built by a team from stanford, meta, and polymarket. loved by 6,000+ across the world comment + RT for free access hedwigmail.com
Made with AI
255
219
1,203
363,315
DFKM!!! 2 whale buys in secs 👀 @mirza is H.I.M man single-handedly took mcap +250k higher $INJ is undoubtedly sprouting 🌱 $1m mcap for sprout before today runs out?
6
1
19
1,161
Endy retweeted
according to what i gathered from my frands, @Hetzner_Online is the best VPS provider out there unfortunately i couldn’t get the $10 monthly plan any other great alternative that doesn’t shutdown, exp server halts and glitches that can drive you crazy? please recommend 😔
I will never use tencent lighthouse VPS anymore In one day the shutdowns alone, server halts and glitches can drive you crazy Like every time my gateway is restarting Someone please recommend a better VPS 😔
1
6
188
Endy retweeted
Sprout it up for me 🌱 $3m mcap atleast 👀
The wait is over. Sprout is LIVE on trysprout.fun 🌱 CA: 0x503f2A7bAc1E2ff2800Aa92EdB82abeF0bFA631E
9
2
24
675
Endy retweeted
1 hour until $SPROUT goes live on Injective 🌱 y’all ready, $INJ community? 👀 also… what are we calling for mcap at launch? if you still need the details on how to participate or what you’ll need for launch, check the quoted post below 👇
Getting ready for @sproutsomefun ? 🌱 Sprout launches tonight. Here's how it works, for anyone planning to take part. Wallets: Keplr or M*taMask are recommended. Connecting to the site will add Injective EVM for you. Bridging: Sprout runs on Injective. Funds on another chain have to be bridged over first, and bridges can take a while. There's a "Bridge to Injective" link in the menu on trysprout.fun. Paired with $INJ: The token will be paired in INJ, so trades are made in INJ. Gas on Injective is paid in INJ too. There is no need to wrap your INJ. Where to find it: Two places only: the Featured page on trysprout.fun, and this account, @sproutsomefun. The contract address gets posted here at launch. Impersonating tokens: Anyone can launch a token called SPROUT, and people likely will. If it isn't Featured on trysprout.fun and the address doesn't match the one posted by @sproutsomefun, it isn't ours. Please tread carefully. Tonight · 10 PM ET · only on trysprout.fun This is information about how the launch works, not financial advice. Do your own research.
1
1
13
671
Endy retweeted
thank you @runupdotfun $RUNNER broke my target of $1m mcap -- did $1.63m founding round participants that sold atp made 4figs w launch on @injective 🤝 anyways might rebuy and watch for a $10m mcap $INJ💙
runup is live runup.fun launch tokens powered by perps. pick a market. long or short. run a token. run it up.
7
3
22
1,302
i shouldn't have touched my banked reset a few hours earlier 😭 ty tibo
Reset all propagated. Enjoy.
1
5
249
Endy retweeted
I built Maro, a web app for a real furniture studio in Lagos called Maro Atelier. They do bespoke design, fabrication and restoration, and like most studios here, their whole intake lives in WhatsApp threads. A customer sends a voice note and some photos, the studio squints at the messages, asks the same three questions five times, and hopes nothing gets lost. I wanted to fix that. So the idea is simple. A customer's messy inquiry, photos, budget, vibes and all, becomes a real project. The app reads the message, pulls out what it actually knows, and then asks for what's missing. It never guesses. If the customer didn't say the wall size, the app doesn't invent one, it just asks. That rule is baked into everything: the AI only works with words the customer actually wrote, and the studio's own rules decide what happens next, not the AI. Once a project has everything it needs, it unlocks booking. Real time slots, Lagos time, real conflict prevention so nobody gets double booked. When a customer books, a booking fee appears on their page with the studio's bank details and their own reference, they pay by transfer, tell the app they've paid, and the owner confirms it from the inbox. Invoices work the same way, any amount the studio chooses. There are two worlds in the app. Customers get a portal with their own sign in, where they see their projects, appointments, messages, quotes and payments, and it updates live as the studio works. The studio gets a workspace of its own, with an inbox, a calendar, invoices, a team section where each staff member gets their own login and only sees their own calendar and projects, and a client dashboard for sending progress updates and quotes. One more thing I care about: honesty. The Our Work page only shows finished projects the customer agreed to share, and it tells the story from the studio's actual records, request to delivery. If a story is an example, it says so. Nothing is invented to look good. I built it because I kept watching good studios lose good customers to scattered messages. Maro turns the chaos into a process the customer can actually see. contra.com/community/ojoU1Za…
8
7
25
640
Obra superpowers is G.O.A.Ted asf
Rough tier list of free github skills i'd actually install in 2026
1
4
194
Endy retweeted
i just successfully entered the @runupdotfun founding sale and secured my $RUNNER tokens on $INJ $1m mcap possible? 👀
$runner whitelist phase is now open run it up runup.fun/coin/0xc4dC11d5866… once it ends in 1 hour, the public curve opens + runup platform goes live. next post will be when platform is live
3
1
17
820
Endy retweeted
I just got my neatHack boarding pass See y’all in 7 days
your neatHack boarding pass is waiting. register asap and grab yours 👀 just 7 more days to go!
4
2
10
635