Founder & CEO @ Volantis - we’re solving the memory bottleneck by using optics to connect huge amounts of fast memory to chips.

SF Bay Area, CA
Pinned Tweet
Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: volantissemi.ai/news-insight…
198
181
1,403
973,202
Tapa Ghosh retweeted
dare i say fidget hardware is the best piece of hardware i’ve ever seen
46
27
505
90,971
Tapa Ghosh retweeted
We've been following @semiDL and team for over a year and are thrilled to be a part of their Series A. Tapa's assembled a world-class team tackling the core problems of memory capacity and bandwidth that hold back model intelligence. Excited to see this team build!
Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: volantissemi.ai/news-insight…
1
1
11
649
Tapa Ghosh retweeted
Proud to be an angel investor in @semiDL and Volantis! We as an industry have a lot of work to do to keep scaling intelligence, and we need all the creative minds we can throw at the problem! Moving data around is a central problem and these guys are working on something truly breakthrough.
Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: volantissemi.ai/news-insight…
2
3
53
7,369
Tapa Ghosh retweeted
Well will soon look back on the days of 50 tok/s as dial-up speeds
Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: volantissemi.ai/news-insight…
2
1
9
549
Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: volantissemi.ai/news-insight…
198
181
1,403
973,202
Insane place to live
SF is so gorgeous Undoubtedly one of the best cities in the world
3
1
39
4,708
SF is so gorgeous Undoubtedly one of the best cities in the world
5
80
7,363
Someone should design a scalable hardware system that brings SRAM-system speeds to large, future models! The world can do much, much better than 300 tok/s/user
ALERT 🚨🚨OpenAI's latest GPT6.1 Sol Ultrafast is NOT running on Cerebras but is instead running at a low batch size on NVIDIA GPUs. What does this say about Cerebras? Will Cerebras be serving GPT6.1 Sol Ultrafast in the future?
3
3
12
3,836
Slower than TPU?
When will AI personal assistants be fast enough to be useful? Using Qwen 3.8 27B on Cerebras at ~1,500 tokens/sec, we made an AI personal assistant 19x faster than a suite of other AI personal assistants - Grok Bot, Meta Muse, and Claude Cowork - on the same dinner reservation task.
1
1
7
1,922
Be different, stand out: use Grok to make your decks instead
Roughly half of the decks I receive these days are Claude-generated, with ~default settings.
1
9
1,108
Today this is unusable in practice because it runs 24X slower than real time New computing hardware will fix that, enabling real time frontier model inference
What can Astra do when given a humanoid embodiment? We built HomeBody to find out. Controlled by GPT Astra, it carries out long-horizon tasks in a previously unseen kitchen—from tidying up across the room to retrieving remembered objects from ambiguous requests—without environment-specific training data or additional policy learning. Here's how we did it 👀: tml.stanford.edu/homebody/
1
1
11
1,598
The impact of ultrafast inference for big models at a low cost is going to be titanic
at this point, what we really need is more tok/sec.
4
2
36
3,339
Tapa Ghosh retweeted
Transformers can learn to predict sequences of natural data without access to any data because they can be trained to approximate Solomonoff induction, the ideal next-token predictor. This is the basis of our project described in x.lingyaoai.com/michaelyli_/status/210… . Here,I explain why you shouldn’t be shocked that this is possible. A really fun project co-led with @michaelyli_ and @KfirDolev and great co-authors @gbruno_dl @ANourya @noahdgoodman, and @YoavLevine
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
2
13
150
22,717
Very cool work
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
6
909
Tapa Ghosh retweeted
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
84
459
3,549
737,143
If you’d rather have 200+ HBM per package in 2027, consider Volantis :)
Replying to @IanCutress
Interposers 200 interposer tapeouts since 2.5D introduced 2028 will have 16 HBMs per package. Mem packaged faster than the standards: HBM3e enabled at 10 Gbps, industry is 7.2-8 HBM4e enabled at 16 Gbps, industry will be 8-10 Gbps
1
2
5
1,263
If you’d rather reach 10,000 GB per chip by 2027, join us at Volantis
Replying to @IanCutress
Memory density scaling to 1600 GB per SoC by 2030. 1200 GB in 2029 1000 GB in 2028 600 GB in 2027 400 GB in 2026
1
1
19
1,445
If only we could use light to solve the memory bottleneck for the entire AI industry :)
We compute with electricity. Why not with light? Light has many advantages for data transmission and perhaps even computation. The semiconductor industry is vigorously pursuing new types of optical computing components for AI. Read the new @bismarckanlys Brief (link below):
1
18
1,371
Prediction: As the bottleneck moves to the physical layer for AI, 1st principles simulators will become more and more important, increasing demand for HPC-like compute systems
4
434