
Getting started with Local LLMs for Software Development
Learn how to use local LLMs for software development with predictable costs, privacy and independence, and the quality required for real-world projects.
Insights, tutorials, and best practices from our team — practical knowledge you can apply today.

Learn how to use local LLMs for software development with predictable costs, privacy and independence, and the quality required for real-world projects.

An async standup bot for Slack built with CopilotKit Channels, Mastra, and a local Gemma model — local-first, typed tools, SQLite as the source of truth.


Agentic, generative UI for Angular — a maintained, MIT-licensed package, shown with a small Angular + NgRx demo.

Retrieval-Augmented Generation entirely client-side: with Angular, Transformers.js, WebGPU, and Google's new on-device model Gemma 4 E2B

Build a local AI coding agent from scratch — Gemma 4 on llama.cpp, three tools, one loop — then learn why running it unsandboxed is dangerous and how NVIDIA OpenShell contains it.

Learn what an Agent Harness is, why it matters, and how it improves LLMs through tools, context management, agent runtimes, guardrails, and intent alignment.

Benchmarking 14 local LLM configurations on Google's A2UI generative-UI protocol on a DGX Spark. The cliff between what works and what fails is sharper than expected.

A deep dive into how AI models really work, from pre-training and inference to Chain of Thought, agents, and the emerging field of Harness Engineering.