If you think you are a good software engineer, just code these beasts. No vibe coding. Do it by yourself.
1
4
191
Loreto Parisi retweeted
Alert: Apple just dropped a new model on Hugging Face. It's a Qwen3.5-9B finetune that turns long documents into small page images to save tokens, then pulls up the full text of only the pages relevant to your question 💡 huggingface.co/apple/LensVLM…
75
278
3,462
281,564
Loreto Parisi retweeted
We've made some small changes to the box, so today is shooting day. And we're literally cooking since the kitchen has the best light in the office office lol Coming soon on lucebox.com
10
6
87
9,030
Loreto Parisi retweeted
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
776
3,495
32,435
10,828,674
Loreto Parisi retweeted
Alibaba’s Zvec team open-sourced zg, a local search tool for developers and AI agents. • Local-first • Works out of the box with popular agents • Semantic, BM25, hybrid, and rg search in one tool Why they built it and how it works:
Article

From rg to zg: Local Search Beyond Keywords

Summary: The information humans and agents need is often scattered across large numbers of local files, making it difficult to locate accurately and efficiently. zg (zvec-grep) is local-first search

86
317
2,782
249,348
Loreto Parisi retweeted
Running Gemma 4 26B A4B on a Mac just got 2x faster! The developer community has been grinding on the mlx.fast leaderboard, pushing Apple Silicon to its limits.
41
150
1,871
148,487
Loreto Parisi retweeted
MLX: Tom Riddle's diary running locally on an iPad powered by Apple MLX with Gemma 4 E4B behind the scene. Children can go crazy for something like this, no? 🔥
18
32
229
32,469
Loreto Parisi retweeted
Follow Ettore if you love Local AI, I bet you'll see incredible creations from him in the upcoming months!
this week was really crazy in AI. And I have only a DGX box for this, not sure how long it will last. What I'm doing in parallel for vllm.cpp: - Qwen 3.8 27b benchmarks (I'd like to take some numbers to show you) - Qwen 3.8 2.4t fitting on a single Nvidia DGX box (yes, we will have it, even if slow) - @Lightricks LTX2.5 support - Index-tts support - @MiniMax_AI Music 3 - @NVIDIAAI 's new Nemotron release - Dspark just landed - dots3-preview support in vllm.cpp And since I'm having issues in overbooking the box with multiple agents, I'm doing a resource-controller api in parallel to control this..
7
2
133
51,767
Loreto Parisi retweeted
Gemma 4 running on an iPhone with just ~500 MB of RAM! User Antikythera on r/LLMDevs shared their calibration-aware quantization approach to shrink Gemma 4 keeping speed and accuracy on an iPhone. They demoed it using an offline assistant that manages calendar actions with just ~516 MB of active RAM.
46
150
2,081
118,280
Loreto Parisi retweeted
Running Hermes Agent locally, powered by DwarfStare with DeepSeek V4 Flash mxfp4 with 1M context is something I'd never imagine possible last year. Now I can't imagine what 2027 will bring to us.
7
3
88
4,868
Loreto Parisi retweeted
DeepSeek V4 Flash 0731 is now open weights! @deepseek_ai has just released the weights for its new flash tier model, DeepSeek V4 Flash 0731. With a score of 50 on the Artificial Analysis Intelligence Index, it lands among the top 3 open weights models on the leaderboard. The weights are released under the MIT license, allowing unrestricted commercial use and modification. DeepSeek V4 Flash 0731 shares identical architecture and pricing with the earlier DeepSeek V4 Flash. At a size of 284B total parameters (13B active), released in mixed FP4/FP8 precision at ~167GB total file size, it lands on our Pareto frontier for Intelligence Index vs. Total Parameters. Among open weights models, DeepSeek V4 Flash 0731 delivers a significant leap in intelligence for its size class. DeepSeek V4 Flash 0731 is also available now through DeepSeek's first-party API. Check out Artificial Analysis to compare DeepSeek V4 Flash 0731 with other leading open weights and proprietary models: artificialanalysis.ai/models
38
88
1,033
67,871
Loreto Parisi retweeted
I'm starting the conversion work from new DeepSeek v4 Flash checkpoint to GGUF. If the model is as good as it looks, I'll probably remove the GGLM 5.2 support from the system, since now we have a model that is best suited for local inference that is smaller. Feedbacks?
82
30
908
63,624
Starlink mini 2026 vs. Starlink 2024, same location, almost same time!
1
1
121
😵😵😵
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!
98
Loreto Parisi retweeted
We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is a breakdown of what’s being fixed and updated in this release: 🧵👇
132
322
3,312
720,755
Loreto Parisi retweeted
llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort! github.com/ggml-org/llama.cp…
17
47
401
85,341
Loreto Parisi retweeted
Local AI just got faster. ⚡️ We worked with @ggerganov to add DFlash support in llama.cpp delivering ~2x faster inference.
llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort! github.com/ggml-org/llama.cp…
18
64
610
104,352