Got a lot of attention for this (thank you!), but also there're some criticisms and questions:
1) it's a stupid idea
Maybe. Again, this is experimental and not meant to be used in production or as a replacement for anything. It is to explore new possibilities.
2) it's unnecessary and sometimes wrong
Existing solutions are not always more correct just because they're "algorithmic". For example pic 1 is Highlight.js and 2 is gpu-lexer, lang is JS not something rare. There are counterexamples too.
Furthermore, language-agnostic is the real reason for exploring this approach. During training, the neural network learned patterns like "[keyword] variable = value" that are not tied to any specific programming language such as:
var ... = ...
let ... = ...
const ... = ...
using ... = ...
val ... = ...
auto ... = ...
type ... = ...
def ... = ...
And that knowledge can be shared without re-implementing a grammar rule repeatedly.
Another advantage is that this approach can potentially generalize to new programming languages or dialects without requiring extensive manual rule creation. If you go to
prismjs.com/test, you can't find popular Web framework languages like Vue, Svelte, Astro, and many others. Not to mention that new languages and dialects are constantly emerging, making it impractical to maintain comprehensive grammar rules for all of them.
When that's a concern, you either use more comprehensive solutions like Shiki with a cost (carefully bundle, detect and load language definitions on demand, larger size), or use an easier approach for all kinds of source code with another cost (less accurate).
That is definitely a good reason to explore this direction.
I trained a small model to do syntax highlighting in the browser with GPU.
Meet gpu-lexer from Vercel Labs: Small (27.5KB), fast (runs on WebGPU), and language-agnostic (model guesses the syntax).
gpu-lexer.vercel.app
It is experimental and built for learning!