CTO @ General Compute

San Francisco, CA
We raised $15m to build the ASICs-first inference cloud. We're betting big on alternatives to GPUs, and the result is that we are already 5-8x faster on most models. Read more about General Compute on Tech Crunch! @FPuklowski @fastinference techcrunch.com/2026/05/28/ha…
30
18
194
34,745
people acting surprised when they learn ultrafast is ultravaluable. it's the same lesson every 6 months
2
263
Jason Goodison retweeted
SemiAnalysis numbers imply OpenAI is selling Cerebras-powered Ultrafast inference at ~$200M per megawatt per year. People are re-running the maths hoping for a typo. We're not surprised. A megawatt is a megawatt. What it earns depends on the silicon inside it. Here's how we think about it: Near term, revenue per megawatt wins. Power is the scarcest input in AI, and whoever monetizes each MW best gets the next MW. Long term, gross margin per megawatt wins. Prices compress, competitors arrive, and only the operators with real cost and pricing advantages keep the spread. Cerebras is built for both. Speed is the product. A frontier model at 750 tokens per second is something GPU clouds haven't publicly matched, and customers like Jane Street pay a premium for it. Premium pricing on a contracted, scarce resource is how you protect margin when everyone else is racing to the bottom. That's why we're extremely bullish on Cerebras, and why we're buying as many megawatts as we can. Premium inference unlocks the most delightful user experience in AI. No more waiting, waiting, waiting, staring at little thinking animations while the model works in the background. Answers that keep up with your thoughts. Our customers are even more excited than we are. That's General Compute.
16
29
163
37,194
Opencode is pretty awesome
7
421
Jason Goodison retweeted
Congrats @FPuklowski and @general_compute team!!!
🚨 BREAKING: Super excited to announce we're deploying the world's fastest inference with @Cerebras. Talk to any developer and they're excited to build with 20x faster AI... the problem is there's almost no compute available. We're here to solve that. Using GPUs for prefill - it's now more affordable than ever too. Thank you to the whole Cerebras team and excited to grow this partnership. First tokens live Q127 🚀

ALT gc-cerebras-640x640-v2-under8mb

2
2
35
5,107
AI music is going to have a big moment in 2028
5
298
Jason Goodison retweeted
@general_compute is buying and deploying @cerebras systems at scale under a multi-year agreement. More builders will now get the fastest inference in the world. In AI, speed is productivity. There is no upper bound to how much faster you want to be. It's true in coding. It's true in agentic flows. General Compute lets inference providers offer that speed without owning the hardware. Cerebras inference will be available through General Compute starting Q1 2027. Excited to work with @FPuklowski and the team.
9
8
121
6,325
Jason Goodison retweeted
Cerebras 🤝 General Compute
🚨 BREAKING: Super excited to announce we're deploying the world's fastest inference with @Cerebras. Talk to any developer and they're excited to build with 20x faster AI... the problem is there's almost no compute available. We're here to solve that. Using GPUs for prefill - it's now more affordable than ever too. Thank you to the whole Cerebras team and excited to grow this partnership. First tokens live Q127 🚀

ALT gc-cerebras-640x640-v2-under8mb

2
2
40
3,224
Jason Goodison retweeted
The world’s fastest inference is coming to @general_compute
12
27
283
45,705
Jason Goodison retweeted
🚨 BREAKING: Super excited to announce we're deploying the world's fastest inference with @Cerebras. Talk to any developer and they're excited to build with 20x faster AI... the problem is there's almost no compute available. We're here to solve that. Using GPUs for prefill - it's now more affordable than ever too. Thank you to the whole Cerebras team and excited to grow this partnership. First tokens live Q127 🚀

ALT gc-cerebras-640x640-v2-under8mb

67
131
1,033
783,982
What I'm most excited about is running Cerebras alongside GPUs. Each phase of inference goes to the chip that serves it best, and all the user sees is ultrafast tokens. A year of work from the team to get here. Proud of everyone who made it happen.
🚨 BREAKING: Super excited to announce we're deploying the world's fastest inference with @Cerebras. Talk to any developer and they're excited to build with 20x faster AI... the problem is there's almost no compute available. We're here to solve that. Using Nvidia GPUs for prefill - it's now more affordable than ever too. Thank you to the whole Cerebras team and excited to grow this partnership. First tokens live Q127 🚀

ALT gc-cerebras-640x640-v2-under8mb

2
12
123
CUDA used to be a real reason not to consider alternative silicon. A new chip meant rebuilding years of software before your first token. This year AI coding started getting you a workable compiler fast, and the chip companies aren't treating software as an afterthought. @SambaNovaAI, @dMatrix_AI, and @positron_ai all treat vLLM and SGLang integration as life or death. What's left is bring-up: getting models running on every chip, the day they drop.
2
269
When I was in the W23 batch nearly everything around me was software. The conventional advice was building hardware was a great way to spend two years delaying finding out if: i) anyone wants what you're building and, ii) if it actually works. The cost of designing hardware has fallen significantly, thanks to software. Quicker and better software unlocks possibilities in hardware. Glad to see YC at the forefront of it. x.lingyaoai.com/tbpn/status/2098135496…
YC's @garrytan says hard tech startups make up a quarter of the newest 200-company Demo Day batch, up 40x since their low point ~4 years ago: "It's 200 companies presenting today. About a quarter of them are hard tech. That's up 40x since our low bar, maybe 3-5 years ago."
5
511
In the past I made multiple youtube video's called "If you build this I will hire you" Essentially an open forum for people to do a take home assignment and add their own flare to it Results were mostly AI slop but every once in a while you get a gem from it. Scrappy, smart kids with vision and work ethic. Seems to work better than the classic resume + leetcode interview style
12
517
We ran some stats on a GPU Cloud vs General Compute Both on standalone requests and an IRL agent coding session The TLDR: @general_compute is much much faster on both This was using @sambanovaAI's last gen SN40 hardware, so just wait to see how fast the SN50 is Here's the breakdown based on harness and build Standalone: TTFT 8,084ms vs 1,136ms -> 7.1x faster Decode 121 tok/s vs 1,025 -> 8.5x faster One request end to end, 34.3s vs 2.1s -> 16.3x faster The bigger bottleneck at some point for long horizon agents becomes CPU operations and storage access. Which colocating with the accelerators will help solve
1
2
15
863
Speculative decoding is how most production inference already runs The question i get from clients isn't whether it works It's whether it works on ASICs It does, and the gains are bigger than they are on gpus The question usually carries an assumption underneath it: that software is what keeps gpus ahead, so better software eventually removes the case for dedicated inference silicon But the same technique runs on the ASIC, on hardware already built for decode The gap widens, it doesn't close x.lingyaoai.com/thejessezhang/status/2…
One cool technique to improve latency in LLMs is to have a small, fast model generate tokens and a larger, slower model review the output in one pass. It'd be like having a junior employee write a report, and then having a senior employee review it in one go before it gets sent to a client. This is called "speculative decoding" and lets you generate faster while maintaining the output quality. Great write-up from the @DecagonAI team on how we've been doing this:
1
1
5
514
What you can build is going to depend an insane amount on TPOT Can't tell you how important this will be for high speed applications
What changes when you're building software for agents not for humans? “The experience is going to be defined by the slowest point in that chain. The number one requirement for agentic query patterns is low latency, because they are executing dozens of SQL queries simultaneously across all these different systems. The most important requirement is the unpredictability of those query patterns, the responsiveness, and the fact that they are much more exploratory than a traditional human report or query.” @ceo_clickhouse Love to hear your thoughts @glcst @kiwicopple @bernhardsson @ashashutosh
2
425
Before General Compute I ran a voice agent company. We had customers in Mexico - with calls in Spanish. And I don't speak a lot of Spanish! Instead, I read transcripts and every pronunciation failure was invisible to me: looked right, but sounded wrong when listened to by a native speaker. Wish I had this benchmark when I was building - but glad it's now out there for others :) x.lingyaoai.com/ArtificialAnlys/status…
Announcing the Artificial Analysis Pronunciation Robustness benchmark, measuring how reliably Text to Speech models say challenging text correctly - Google Gemini 3.1 Flash TTS leads at 88.1%, followed closely by SpaceXAI TTS at 87.6% and ElevenLabs Eleven v3 at 85.6% Existing Text to Speech (TTS) evaluations, including our TTS Arena, capture overall listener preference (e.g., how natural a voice sounds), and Word Error Rate (WER) checks whether the right words come out. Pronunciation Robustness adds a view of whether those words are said correctly (e.g., reading “St.” as “Saint” and “Street” in “St. Mary’s is on Church St.”) - this is critical for production voice agents, which need to get account details, names, currency amounts and more right to be trusted by users, at the low latency that conversational experiences demand. Overview of Pronunciation Robustness Each model generates audio for 454 sentences containing 701 target words or phrases, across four categories: 1. Contextually appropriate (words read differently depending on context, e.g., a wound that is bandaged vs. a bandage that is wound) 2. Expanding shorthand (numbers, dates, units and notation read out naturally, e.g., 6'2", Chapter XVII or 1 tsp of sugar) 3. Preserving exact sequences (codes, paths, emails and identifiers spoken exactly, e.g., .env.local or a.chen@ucsf.edu) 4. Standalone terms (brand, place and technical names, e.g., Arkansas, façade or genre) Sentences are sent as written, with no normalization on our side beyond each model's default setting. Screened human listeners judge whether each target was pronounced correctly against accepted pronunciations set in advance, excluding anyone who fails an attention check. The score is the share of correct judgements, excluding "I could not tell" answers, and we publish a model once 95% of targets have three or more approved listeners. Key results: ➤ Overall leaders: @GoogleDeepMind’s Gemini 3.1 Flash TTS leads at 88.1%, followed by @SpaceXAI’s TTS at 87.6%, @ElevenLabs’s Eleven v3 at 85.6%, v3 Conversational at 84.8% and @Alibaba’s Qwen-Audio-3.0-TTS-Plus at 81.6%. Gemini 3.1 Flash TTS leads Contextually appropriate and Expanding shorthand, SpaceXAI TTS leads Preserving exact sequences, and Qwen-Audio-3.0-TTS-Plus leads Standalone terms. ➤ Hardest categories: Expanding shorthand (62.4%) and Preserving exact sequences (62.9%) trail Contextually appropriate (86.2%) and Standalone terms (86.1%). We expect Expanding shorthand and Preserving exact sequences scores to rise considerably with normalized input text, especially for models without normalization on by default. ➤ Preference ≠ pronunciation robustness: The most preferred voices aren't always the most accurate - Sonic 3.6 ranks #1 on the Provider Voice Arena at 1276 Elo but #11 on pronunciation robustness at 74.5%, while Gemini 3.1 Flash TTS ranks #9 on the Arena at 1201 Elo and #1 on pronunciation robustness at 88.1%. See more details below ⬇️
3
12
829
My girlfriend and I are moving into a new apartment. Fable built us a full 3d model of how the furniture will fit in < 10 mins
1
6
384
Flow state is still alive in software engineering What that looks like now vs 12 months ago has changed massively though, at least for me Before it was writing lines of code against a spec (mostly unaided by AI) Now it's managing agent sessions, context switching, and input -> output has to be sub 1 minute for me to be in flow Genuinely the only reason I can still reach flow now is by using our own inference cloud
2
8
445
Jason Goodison retweeted
Full house at the @general_compute + @SambaNovaAI + @pipecat_ai voice agents hackathon at @AGIHouseSF. Thank you also to the @GradiumAI team for supporting with speech-to-text and text-to-speech APIs.
8
7
40
2,565