MTS @ context.dev classifier.dev. Scout @trygravityai & investor. Most startup problems can be solved w/ 100M views.

San Francisco
classifier.dev now outperforms jev and is free go nuts guys
152
354
5,868
499,802
is there a way to show improvement on long-horizon coding tasks without actual frontier model rollouts? for example, I want to benchmark @getcontextdev's highlights model to see how much it saves on long-horizon coding tasks and measure its impact on accuracy/performance but I don't want to spend 50k of API credits! and if there's a model change, I don't want to spend ANOTHER 50k IN CREDITS. anyone found a solution to this? would help a lot
Prediction: universities will get priced out of AI evals work We just spent $50k yesterday benchmarking 40 long horizon coding tasks As we design harder tasks for agents, trajectories are getting longer and need more tokens. That means substantially more inference and compute. How is this economically feasible for academics working off university grants?
2
3
448
POV: you prompted poorly and accidentally sent your boss hundreds of emails instead of 5 am I fired?
9
62
4,419
pre-made training apis are a bit antiquated in the era of autoresearch constraining llms makes zero sense now. give them the raw gpus and they'll figure it out even tinker is a bit constrained (although mega-convenient) let the llms free. give them the ability to spin up 8/12/more GPUs on zero notice let them choose the base model. let them mess around and find out
Replying to @michael_chomsky
Interesting and thanks for sharing . Would love to understand if the need to ssh is because you are doing some custom algo? Or models not easily available to train or something else?
3
1
19
1,775
if you like this you'll love garlic.ai frontier models, e2e encrypted, free forever you don't need to use subpar local models for private ai. with confidential computing, literally nobody ever reads your packet. garlic doesn't see it, the downstream provider doesn't see it it's encrypted before it leaves your browser
Anyone can get my friend’s real SSN from Claude I'm Terrified of my data in training sets. Emails in databases. iMessages harvested. Cant trust the cloud as AI is crazy good at hacking So I built Underdog for myself: on-device AI OS. Capable & 100% local. now my friends love it
1
9
1,103
this is so dumb token prices are about to fall 50%
inference margins are just getting started if you own electron to token, benchmark levels are 50% margins on blackwells, set to increase to 90% (good lord) on feynman inference businesses are exploding and it's really still so early
2
1
24
3,484
Michael retweeted
Training custom models is both simpler and much more annoying than I expected We have some fun stuff cooking up!
8
2
42
1,558
why is everyone renting out GPUs by the hour? my training run takes 20 minutes on 8 GPUs. i want to be charged by the minute modal is ahead of everyone else here, the dx is insane. unfortunately pretty expensive.
27
2
242
17,317
i don’t completely understand the gpu/datacenter/powered land market but there seems to be a massive trust/information problem buyers need to know if the supplier can actually fill and if they’re shady/bad actors/scammy/etc and their actual timeline sellers need to know creditworthiness/funding timeline who’s solving this?
18
3
68
4,177
“hey guys anybody gave 2GW laying around? budget is 2B/yr for electricity (flexible) pls dm if you have some!” wtf is going on
Looking for 500MW (minimum 250, max 2GW) of powered land for a PortCo. Company is completely price insensitive and has billions behind it. Ideally electrified shell is already built. I understand everyone wants this, but figured I’d throw it out there!
20
15
1,052
101,560
ok since this tweet i have become a compute/colo broker send me your colo/powered land/gpu details and i'll match you with a buyer this will be fun!
33
4,175
ngl jev on edge gpus is kinda op optimized 27b models can run in 30ms so the over-the-wire time can be longer very cool ship by cf
introducing 𝚌𝚕𝚎𝚏: our first models trained by @cloudflare's workers ai team. today, we're releasing two fast and accurate decision models that top the benchmarks for quality and latency. use them hosted on workers ai or grab the weights from @huggingface, because we open-sourced it too. blog.cloudflare.com/clef-dec…
2
75
5,054
algorithm loves me this week, everything I post is going viral to prove it, here's a picture of my pet: U+1F431
4
38
2,161
cool insight let’s say i tell an llm to insert a rare chinese token wherever there should be a segmentation (chunk) boundary we can get the probabilities of the segment at all token locations, but each probability is based on all previous tokens not being boundaries! curious if anyone can think of a solution to this. for example pretty sure a diffusion model can predict all possible segmentation tokens at once (maybe?), but not sure how to get past this with a traditional decoder! food for thought
Replying to @michael_chomsky
The main issue I see here is that the LLM doesn't see any earlier chunk boundaries when deciding later ones. If your text has 10 chapters, the LLM only places a chunk boundary after chapter 7 if it thinks that despite 1-7 belonging together, 8 will deserve a new chunk.
2
7
1,103
why is some random dude in a 4.5M seed announcement lmaooo congrats photon, happy customer!
We raised a $4.5M seed to let agents text people. Since April: 50K+ developers, 450K monthly npm downloads and 10x revenue. Bring your agent to iMessage, WhatsApp and more → tryphoton.ai
5
35
3,718
introducing devreljobs 2.0 we now use: @getcontextdev for high-resolution icons @getcontextdev to find more job opportunities @getcontextdev to enrich job opportunities if you're a devrel, and need a job, go to devreljobs then dm me because I match devrels with top companies the first question I ask people? what companies on devreljobs do you want to work for? find the employment of your dreams now at devreljob.com
3
3
42
1,959
me to codex: "train faster" codex: "ok I guess I'll just rent a zillion gpus" codex realizes there's no way to get 8 gpus on zero notice anybody have a simple tool where i can 8-30 high-end GPUs on zero notice?
8
34
3,735
if anyone wants something shipped for context, just dm me and there's a good chance I get it out in <24 hours
Replying to @mynameisyahia
You’re solving my use cases faster than I can try using Context for them
3
17
1,240
btw this is the same trick that makes spec decoding work if we have a small model propose the next n tokens, we can check them in one forward pass! as long as the small model is calibrated (usually in agreement) with the larger one, you should see an increase in inference speed because you save on expensive large-model forward passes!
incredibly clever chunking trick by the zeroentropy team (now at notion) llms are autoregressive, which means generating text is super slow. but passing text in and finding the logprobs of a particular token at all locations takes a single forward pass! so we can ask the llm, 'like yo llm, please chunk this by putting this unlikely character 段 wherever there should be chunks" and we instantly get the segmentation locations (well, not instant, we still need a forward pass which means we need to prefill) but the advantage is super accurate chunks in the time it takes to prefill!
1
9
1,025
incredibly clever chunking trick by the zeroentropy team (now at notion) llms are autoregressive, which means generating text is super slow. but passing text in and finding the logprobs of a particular token at all locations takes a single forward pass! so we can ask the llm, 'like yo llm, please chunk this by putting this unlikely character 段 wherever there should be chunks" and we instantly get the segmentation locations (well, not instant, we still need a forward pass which means we need to prefill) but the advantage is super accurate chunks in the time it takes to prefill!
14
24
533
33,937