gluecks retweeted
now that programming is easier than ever can we leave Electron and go back to native desktop apps?
141
241
6,676
179,711
gluecks retweeted
Replying to @thepirateface
Why isn’t anyone talking about how these guys went out of their way to gate the model, collect emails, and then spam users with advertising?
2
1
20
1,753
gluecks retweeted
You might've of heard video game decompilations using A I tools, but now someone vibe coded a video driver to support for NVIDIA’s GeForce 10 series (aka Pascal) to work in Windows XP, based off NVIDIA 368.81 driver, dubbed as Forceware 382.69. 😭 32-bit: github.com/SupraGSX/Forcewar…
132
320
4,423
268,908
gluecks retweeted
We release Whistle: speech to text in a 16.9MB file, running on the CPU in the same engine as Needle. It mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish. cactuscompute.com/blog/whist…
92
193
3,171
307,562
gluecks retweeted
Introducing Phonon-2: a new standard in speech recognition per byte. At just a tiny 164MB download, it is more accurate on average than OpenAI's Whisper large ( a model 10X its size) Transcribe an hour of audio in just 20 seconds on a MacBook Air! Open weights, CC-BY-4.0, today. @fermion_ai
224
418
6,153
558,826
gluecks retweeted
woah. Court just unsealed Summary Judgment motions in New York Times v OpenAI / Microsoft. Easy to see why the AI companies had so many redactions (pink). Their own people wrote the lede for NYT, the "largest theft of labor in human history"... "more and more substitutive..." /1
60
1,522
4,726
851,024
Replying to @Pirat_Nation
>wanted to find a team he could trust >picks the most untrustworthy people imaginable
13
104
6,221
104,391
Bad news, Nexus Mods has acquired SteamDB, one of the most useful tools for tracking Steam games SteamDB creator xPaw is stepping back after years of running the site largely on his own, he said the growing workload eventually led to burnout, and he wanted to find a team he could trust to take the project forward. The bigger plan is to connect SteamDB’s game version and update data with Nexus Mods.
1,092
1,480
24,874
1,830,984
At the risk of sounding LLMish: A/B testing is okay (I guess). People gotta eat. @ProtonPrivacy claiming it's a bug / caching, out of your control, and even saying it's "a remnant of a sale" when you show users higher prices at random is all beyond WTF. #1 for a VPN: trust.
Why is @ProtonVPN lying about running price sensitivity testing on their customers? Their pricing page gives you one of two prices at random. This is A/B testing, but instead they're saying the price difference is due to remnants of a sale. Their own code says otherwise. 1/5🧵
34
gluecks retweeted
STEAM MACHINE + SSD GIVEAWAY! It's your chance to win a Steam Machine 512GB with Controller, a 1TB PLAY X SSD to store all your games, and a Steam gift card to buy even more 👀 Want more chances to win? There's also a 2ND PLACE prize of a 1TB PLAY X + Steam gift card! To enter: ➕ Follow @Lexar_Gaming  🔁 RT this post 🤝 Tag TWO friends Closes 10th August
3,668
3,729
3,209
151,494
gluecks retweeted
Replying to @BeanJuiceStudio
No the review is correct, the UI in those screenshots screams AI. It has so many tells that are hard to articulate, but if you ask an AI to code you a dashboard its what they look like. It's the glassy look, the rounded coloured buttons, the overall hierarchy of the text. 99% AI
1
1
2
61
gluecks retweeted
how did we make deepseek outperform opus 4.7? i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem. context: spent the two days looking at billions of tokens in @CommandCodeAI (tb open source ai cli) using deepseek. I ended up writing a tool-input repair layer. the trigger was watching deepseek-flash fail on the simplest /review run, every shellCommand and readFile call bouncing back with a raw zod issues blob, the model unable to recover because the error wasn't in a form it could read. by the end deepseek v4 pro was beating opus 4.7 6/10 times on our internal evals. a few things i learned that feel general: 1/ the failure modes aren't random they're a small finite compositional set. across deepseek-flash, deepseek v4 pro, glm, qwen, the same four mistakes repeat almost exactly: - sending `null` for an optional field instead of omitting it - emitting `["a","b"]` as a json *string* instead of an actual array - wrapping a single arg in `{}` where the schema expected an array (an "empty placeholder") - passing a bare string where an array was expected (`"foo"` instead of `["foo"]`) four repairs, ~30-100 lines each, ordered carefully (json-array-parse must run before bare-string-wrap or `'["a","b"]'` becomes `['["a","b"]']`). that is the whole catalogue. when i hear "this open source model can't do tool calls" i now assume one of those four, and so far that's been right ~90% of the time. 2/ the funniest failure mode is also the most revealing. deepseek-flash, when asked to edit or write a file, sometimes emits the path as a *markdown auto-link*: filePath: "/Users/x/proj/[notes.md](http://notes. md)" our writeFile tool obediently trued creating files literally named `[notes.md](http://notes .md)` until we caught it. this is not a hallucination. it's the post-training chat distribution leaking through the tool boundary the model has been rewarded for auto-linking in conversational output, and is applying that prior in a context where it makes no sense. the fix is two regex lines that unwrap only the degenerate case where link text equals url-without-protocol real markdown like `[click](https://x .com)` passes through untouched. this is also conditioning of their own tools during RL which were different from all other tools we write and ofc can't predict. "tool confusion" is a more useful frame than "capability gap." the model knows how to format a path. it just hasn't been told clearly enough that this path is going to fopen, not into a chat bubble. so we encode that hint at the schema level `pathString()` instead of `z.string()` and the leak is plugged for every path field at once. 3/ the design choice that mattered was inverting preprocess-then-validate to validate-then-repair. my first attempt was the obvious one: a preprocessing pass that normalized inputs (strip nulls, parse stringified arrays, etc.) before zod ever saw them. it broke immediately, writeFile content that *happened* to be json-shaped got rewritten before it hit disk. silent corruption, easy to miss in a smoke test. then i made it less greedy - parse the input as-is. if it succeeds, ship it. valid inputs are never touched. - on failure, walk the validator's own issue list. for each issue path, try the four repairs in order until one applies. - parse again. on success, log `tool_input_repaired:${toolName}`. on failure, log `tool_input_invalid:${toolName}` and return a model-readable retry message. the structural insight here is: when you preprocess, you encode a prior about what's broken. when you let the validator complain first, the schema is the prior, and you only spend repair budget at the exact paths the schema actually disagreed at. the validator is doing the work of localizing the bug for you. it's the same shape as cheap-then-careful everywhere else try the fast path, fall back on evidence. (this also gives you per-tool telemetry for free. you can watch repair rates per (model, tool) and notice when a model regresses on a specific contract before users do.) 4/ shape invariants and relational invariants need different fixes. the four repairs above all handle shape problems wrong type, missing key, wrong container. but read_file had a *relational* invariant: "if you provide offset, you must also provide limit, and vice versa." deepseek kept calling `readFile({ absolutePath, limit: 30 })` and getting an `ERROR:` back. you can't fix this with input repair, because each field is independently valid the bug is in the relationship between them. so i taught the function the model's intent instead. `limit` alone → `offset = 0`. `offset` alone → `limit = 2000` (matches common read tool ops default). then surfaced the decision back to the model in the result: "Note: limit was not provided; defaulted to 2000 lines. To read more or fewer lines, retry with both offset and limit." no `Error:` prefix, so the tui doesn't paint it red. the model sees what we picked and can self-correct on the next turn if our guess was wrong. transparency over silent magic wins big. repair where you can. extend semantics where you can't. surface the choice either way. zoom out: a lot of what looks like model capability is actually contract design. a strict schema is a choice with a cost it filters out noise, but it also filters out recoverable noise from any model that hasn't memorized the exact json contract you happened to pick. the largest commercial models eat that cost invisibly and are linient on tool calling because they've seen enough of every contract during pretraining; open models pay it loudly and get dismissed for it. the harness is where you mediate between distributions. four small repairs (i'm sure more to follow as we have three more merging today), two regex lines for auto-links, one relational default, one prefix change. the model didn't change. the contract got more forgiving in exactly the places it needed to be. deepseek v4 pro now beats opus 4.7 6/10 times on our internal evals. imo "skill issue" applies to the harness more often than the model.
Wow I just made DeepSeek V4 Pro beat Opus 4.7 6/10 times in our internal evals by auto repairing many of its quirks in tool calling. It’s performing super solid for such a cheap model.
85
196
1,989
2,270,782
gluecks retweeted
GIVEAWAY I AM GIVING AWAY 1 PC Collector's Edition for 007 First Light and a Limited Edition Controller Just Like, RT, and Comment. The controller will go to someone who replies! Enter here: gleam.io/KkkZr/007-first-lig…
843
1,084
1,395
89,786
I found the "Quote" button ☺ Thanks!
store.steampowered.com/app/3… 4/1 23:59까지 [Dice Eater] 무료 공개! 4/1 23:59까지 이 게시물을 X에 인용하여, [Dice Eater] 유료 구매 내역을 인증할 경우, [Staffer Reborn] 게임키 증정! 4/1 23:59 まで『Dice Eater』無料公開! 4/1 23:59 までにこの投稿をXで引用し、『Dice Eater』の有料購入履歴を認証していただければ、『Staffer Reborn』のゲームキーをプレゼント!
1
79
@Team_Tetrapod What happened to these Steam key rewards?
1
65
To celebrate the launch of Portal Fantasy, we're giving away 2 copies for Steam 🥳🎉To enter, like and share the post, plus tag a friend in the comments! Winners announced on April 17th ❤️A huge thanks to @PortalFantasyio for partnering with us to offer you this opportunity!
48
42
53
2,129
RT @MSportgames: To the NASCAR game community! From December 31, 2024, all NASCAR game titles and their DLC content will no longer be avai…
98
#GIVEAWAY 🥳 We're offering a key of Melobot - A Last Song on Steam for International Games Day! 🎮 To participate: 🤖 RT + FOLLOW 🤖 Invite 2 friends to participate Ends on November 22th! Good luck! 🎶 #indiegame #cosygame #gamedev #music
40
38
45
2,106
📷 #GIVEAWAY TIME! 📷 Want to join the sticker adventure in #ATinyStickerTale by our friends at Ogre Pixel? We're giving away a FREE Steam key! Follow @nelpastelstudio and @OgrePixel Retweet this. Tag 2 friends who love games! Winner will be chosen on October 14🍀📅 #IndieGame
58
50
50
3,997
Time for a 🙌 DROS KEY GIVEAWAY 🙌 We need you to give this Dros a name! Reply with your answer to go in the draw to WIN a Steam Key for Dros when it releases. Make sure you like 🩷 RT 🔄 and follow ☺️ us too.
68
62
105
13,212