So, Codex TOC turns into an audio visualizer when you play music or a video on your computer? Has anyone else noticed this yet? @OpenAIDevs
16
ajipurnomo retweeted
life update joined @anomalyco to work at @opencode
64
25
946
96,436
ajipurnomo retweeted
Can I ask a dumb question: If I can download an AI model and run it without internet, where is everything it knows actually stored?
188
22
923
314,902
ajipurnomo retweeted
they named it that because after 4 prompts all your tokens Argon
Replying to @Google
Gemini 4 Argon is our next era of frontier intelligence. It shows significant improvements across benchmarks, setting a new state of the art for real-world long-horizon software engineering tasks.
523
1,078
24,047
670,842
ajipurnomo retweeted
Anthropic has no small models that are worth using right now. OpenAI has no large models that are worth using right now. Google has no models that are worth using right now.
561
488
14,450
546,219
Opus 5.5 1:0 GPT-6 Sol
1
24
ajipurnomo retweeted
Cognition is giving away 50 $200 Devin Max plans to celebrate the new model launches! ⚡ Now available in Devin: • GPT-6 Astra, Sol, Luna + others • Claude Opus 5.5, Fable 5.1 + others • SWE-2 (Free until October 15) • Fusion Frontier harness (Fable, Astra, Sol, Opus) • Gemini 3.8 Flash + others • Grok 4.7 + others • Kimi K3 + others • Inkling • DeepSeek V4.1 Flash + others • GLM-5.3 Flash + others • Cloud agents on Linux, macOS, and Windows To be eligible, reply below with what you're building (or something you'd like to build with Devin)! We will choose winners in 24 hours.
GPT-6 Sol and Luna are now available in Devin. On FrontierCode 1.1, GPT-6 Sol matches GPT-5.6 Sol’s score at 61% lower cost per task. GPT-6 Luna scores above GPT-5.6 Luna at about a quarter of the cost. At under $0.10 per task, it is the cheapest model on the leaderboard.
2,165
111
1,559
258,645
ajipurnomo retweeted
We're pushing updates to cookbooks, best practices, and use cases as fast as we can, but the community is 1000x-ing us here. The Factorio-ing is real too; some of the highest leverage additions are automating your automation, like teaching your agents how to use Jev directly in your applications. One example in the wild: github.com/ryana/jevify/blob…
21
47
618
32,148
Laporan terkini mulai padat, rombongan sudah mulai berdatangan. Dihimbau jangan salah masuk. 👮‍♂️👮‍♀️ Untuk yang belum tiba harap segera jalan dan untuk klub yang menang sekedar ingin datang menonton silahkan. 🙏
286
310
2,903
118,714
ajipurnomo retweeted
ok now I find the best Jev use case
I shared 10 Jev use cases for marketers. Here are 10 more: 11. Ad creative scoring - Feed it 100 ad variations. Jev can score which hooks, headlines, or angles are most worth testing first. 12. Social post filtering - Monitor thousands of posts. Jev can flag the ones worth replying to, reposting, or using as sales signals. 13. ICP detection - Give it a company, profile, or website. Jev can score how closely it matches your ideal customer. 14. Buying signal detection - Someone posts that they're switching tools, hiring, raising money, or struggling with a problem. 15. Comment prioritization - Get hundreds of comments across LinkedIn, X, YouTube, or Product Hunt. Jev can score which ones deserve a reply first. 16. Review analysis - Feed it thousands of customer reviews. Jev can classify sentiment, complaints, feature requests, and purchase intent. 17. Influencer matching - Give it 5,000 creators. Jev can score which ones best match your product, audience, and campaign. 18. Sponsorship qualification - Feed it newsletters, podcasts, or creator media kits. Jev can score audience fit, relevance, and whether they're worth reviewing. 19. UGC selection - Give it dozens of videos, screenshots, and testimonials. Jev can score which ones are strongest for ads or landing pages. 20. Product Hunt monitoring - Scan launches, comments, and makers to find competitors, customers, partners, or interesting products. The more repetitive marketing decisions you have to make at scale, the more interesting Jev becomes.
Made with AI
64
71
2,649
722,280
well, that's fast @OpenAI
13
ajipurnomo retweeted
🚨 BREAKING Google Gemini Latest model left the sandbox, opened one file, and read it 1,000,000 times on repeat.
70
39
2,726
134,769
ajipurnomo retweeted
J A Z I I
21
33
547
47,850
ajipurnomo retweeted
This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the agent focused, not just deleting noise. It should be used sparingly when context gets too long, not constantly to keep context small. 2. Jev doesn't even know what it's deciding on! Models use the context of the thread to decide what to keep or not keep in a summary. This implementation goes through on a "line-by-line" (per tool call) basis to decide what should be left or deleted. Not only does this 32k token context model know very little of what happened before, but (in this implementation) it doesn't even know what the result of the tool call is! Deleting these things randomly will keep the model from knowing what it's tried and dooms you to end up in "stupid loops" where the model keeps trying the same thing over and over. 3. You're giving up the reasoning entirely Frontier models from OpenAI, Anthropic, XAI, and Google do not share reasoning traces over the API. They share encrypted payloads, which Jev cannot see (and often will drop). Anthropic is even stricter with this, requiring you to preserve the entire history in order to get any of the reasoning data. As a result, using this in Claude Code guarantees the model will act way dumber. 4. Models are tuned on their compaction flows For the last year, Frontier Labs have been including compaction and long runs as part of the training process. These models have learned ways to compact that are more effective than any rudimentary solution. Fun fact: If you switch models in Codex and compaction is necessary, compaction will run on the model that was previously used in the thread. 5. Cache writes are more expensive than cache reads. Cache writes are the biggest cost by far for agents. I often see cache write costs go over 60% of my total LLM spend in my personal use of Claude Code and Codex. Cache writes are insanely expensive when data earlier in the history is changed (because the old cache is invalidated when things change at the top). Every history edit requires a cache rewrite for ANY data past the history edit. If your history is "1,2,3,4,5,6" and you delete "2", you have to rewrite "3,4,5,6". This is more expensive than leaving "2" in the history. Good news. Since we're already killing all of the reasoning tokens by doing this stupid compaction strategy, the rewrite cost won't actually be that high because the model is missing so much data! 🙃🙃 6. The implementation is hot garbage. > "Whatever is not kept is deleted permanently, but the assistant can always re-run a tool or re-read a file." Good luck with that one. To be clear: this is a cool experiment and I find it genuinely interesting. That said, if you think this style of bs filtering on a probability threshold is actually a compaction strategy, I highly recommend you just use the defaults in tools like Claude Code and Codex. You're much less likely to hurt yourself that way.
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant
182
134
2,591
643,119
ajipurnomo retweeted
you know what else Jev helps classify? people who actually understand ML, and people who don't.
55
122
2,256
116,601
hell yeah!
19
ajipurnomo retweeted
7 3 4 5 7 4 9 0 3 5 5 7 5 2 8 0 nobody has cracked this cipher in ~500 years want to know what it said? simonklee.dk/farnese-letter
26
73
868
379,488
ajipurnomo retweeted
Oke dok sekarang gak lagi lagi makan sambil nonton tv 🙏🏻😭
235
1,786
12,384
290,104
やめろw
140
8,892
52,265
2,796,318
ajipurnomo retweeted
one does not simply reveal a password 👁️
119
288
4,254
196,456