Google DeepMind’s new EmbeddingGemma 2 is a small 740 million parameter model made for running directly on everyday devices like phones and laptops. Instead of focusing only on text, it can work with text, code,…
Mistral has released a preview of ML4, internally nicknamed “Le Chonk.” Yes, that is actually what they call it. The model has 1 trillion parameters, although only 49 billion are active at a time. Mistral…
Open WebUI released two updates close together this summer that changed how tools and plugins work in practice. Version 0.11.0 arrived on July 27, 2026, followed by v0.11.1 on August 25. While 0.11.0 also brought…
LattePanda launched Mu Ultra on September 9, 2026. This small x86 compute module comes with either a Core Ultra 5 226V or Core Ultra 7 256V processor and 16GB of memory, designed to run AI…
PewDiePie has revealed Ajax, a local AI model he's building as a personal assistant for his Odysseus workspace. It's designed to run on your own hardware, work with your everyday tools, and refuse fewer requests.…
Aleph Alpha has released Kolibri-1, an open-weight Mixture-of-Experts model for English and German. It has 78.1B total parameters, but only 3.46B of them are active for each token. For a model focused on just two…
There’s a new AI trend where people are merging games, and the results look incredible. Legal and technical issues aside, this article explains how AI is being used to merge games and whether local AI…
Having an AI model is cool, but have you ever wondered what would happen if your local AI could run your script and use the output to solve a problem? Probably you didn't, but if…
Ai2 has released AstaBrief 8B, an open-weights model designed to take research questions and scientific sources and turn them into cited reports. But unlike a normal chatbot, AstaBrief isn't a multipurpose model. It has a…
Cloudflare has released Clef and Clef-flash, its first open-source decision models. Both models are available on Cloudflare Workers AI, while their weights are also available publicly under the Apache 2.0 license, meaning developers can download…
Why can a 16 GB model run on an 8 GB GPU? Quantization. This article shows how it works with a small Python script, then explains the GGUF file format and what the name Q4_K_M…
The whole article is designed to answer the question for people who wonder if VRAM is really essential for AI models, but if you need a quick answer, the answer is yes. Weights AI models…