Lokale LLMs: Sind sie den Aufwand wert? — Rainer Hahnekamp — entwickler Summit, Berlin, 17. September 2026
Unternehmensvorstellung von Soverius AI
Soverius AI

Agenda

LOKALE MODELLE FÜR KI-GESTÜTZTES CODING 01 WARUM Warum lokale Modelle? 02 WIE llama.cpp + Gemma 4 03 DEMO Capstone 04 PROBLEME Geschwindigkeit und Qualität GESCHWINDIGKEIT Hardware QUALITÄT Loop / Graph Engineering 3
Soverius AI

Warum lokale Modelle?

LOKALE MODELLE FÜR KI-GESTÜTZTES CODING 01 WARUM Warum lokale Modelle? 02 WIE llama.cpp + Gemma 4 03 DEMO Capstone 04 PROBLEME Geschwindigkeit und Qualität GESCHWINDIGKEIT Hardware QUALITÄT Loop / Graph Engineering 4
Soverius AI

Semi-autonome Loops

Zoom folgt dem Ablauf
Der gesamte semi-autonome Loop vom Issue bis zum MergeEin Mensch erstellt ein Issue in GitHub oder GitLab. Im Loop wird die Tauglichkeit geprüft und durch Rückfragen mit dem Menschen geklärt. Sobald das Issue bereit ist, gibt ein Programmierer bei Bedarf den Start frei. Ein Coding-Agent implementiert. Performance-, Accessibility- und Architektur-Agents reviewen. Ein koordinierender Agent kann die Reviews steuern, bündelt die Befunde und entscheidet über eine weitere Implementierungsrunde, eine Rückfrage an einen Menschen oder den Abschluss. Die Änderung wird als Pull Request bereitgestellt. Ein Programmierer prüft und bestätigt, dann wird gemerged. Rückfragen & Klärung bereit Weitere Runde fertig IssueMensch · GitHub / GitLabTauglichkeit prüfenIst das Issue umsetzbar?Programmiererggf. Start freigebenImplementierenCoding-AgentReview-AgentsPerformance-AgentAccessibility-AgentArchitektur-AgentKoordinierenderAgentSteuert Reviews · entscheidetMenschbei BedarfPull RequestÄnderung + Review-ErgebnisProgrammiererPrüfen & bestätigenMerge
0:00.0 / 0:36.0

Was passiert, wenn solche Loops auf lokalen Modellen laufen?

5
Soverius AI

Warum lokal?

Gründe für lokale Modelle

€

Kostenkontrolle

Schloss

Datenschutz

Getrennter Stecker

Unabhängigkeit

6
Soverius AI

Setup und Konfiguration

LOKALE MODELLE FÜR KI-GESTÜTZTES CODING 01 WARUM Warum lokale Modelle? 02 WIE llama.cpp + Gemma 4 03 DEMO Capstone 04 PROBLEME Geschwindigkeit und Qualität GESCHWINDIGKEIT Hardware QUALITÄT Loop / Graph Engineering 7
Soverius AI

Einstiegssetup

AGENTClaude Code
ruft auf
INFERENCE ENGINEllama.cpp
führt aus
MODELLGemma 4
8
Soverius AI

Demo: lokale Inferenz

9
Soverius AI

Die Capstone-Aufgabe

LOKALE MODELLE FÜR KI-GESTÜTZTES CODING 01 WARUM Warum lokale Modelle? 02 WIE llama.cpp + Gemma 4 03 DEMO Capstone 04 PROBLEME Geschwindigkeit und Qualität GESCHWINDIGKEIT Hardware QUALITÄT Loop / Graph Engineering 10
Soverius AI

Qualität und Datensouveränität

Geschwindigkeit ist eine Frage des Hardwarebudgets. Leistungsfähige lokale Hardware kann mit Cloud-Inferenz mithalten.

11
Ein leistungsstarker Sportwagen im Stau und vor Baustellen

Der Porsche-Test

Brauchen wir für jede Fahrt die volle Leistung?

Die meisten Softwareaufgaben erreichen nie die Autobahn ohne Tempolimit und Baustelle.

Soverius AI

Mindestvoraussetzungen

Mindestvoraussetzungen statt Benchmark

Mindestanforderung
Modell A
Modell B
Modell C
Kriterien
✓Funktioniert mit Coding-Agents
✓Bewältigt long-running tasks
✓Mindestens 128K Token Kontext
13
Soverius AI

Die Capstone-Aufgabe

Die Full-Stack-Aufgabe

Angular

Angular

  • Feature-Komponente
  • UI-Komponente
  • SignalStore
  • Unit-Test
Spring

Spring

  • JPA-Entities
  • Mapper
  • Controller
  • Datenbankmigration
Playwright

Playwright

  • End-to-End-Test
14
Soverius AI

The Capstone Prompt

/goal

In my holidays overview I want a German vocabulary trainer, similar to the existing quiz feature. For destinations in German-speaking places (currently Lübeck and Vienna) users should be able to practice a few basic German words before they travel: the holiday card shows an icon when a vocabulary test is available, and opening it lets the user translate words.

Do a full-stack implementation (backend + frontend) following the structure and conventions of the existing quiz feature (architecture, feat folder, backend entity/repository/controller). The app must be accessible, and the accessible names below are required. Create and run e2e tests for the feature.

Acceptance Criteria

1 Successful Verification
- The frontend build check is pnpm ng build --optimization false.
- Start any required application dependencies using the project's existing setup, then start the backend with ./gradlew bootRun.
- Wait for the backend to become ready and verify that both the existing /heartbeat and /holiday endpoints respond successfully. Starting the process alone is not sufficient.
- Start the frontend on port 4200 and run the vocabulary feature's e2e tests with the project's Playwright setup.
- If compilation, application startup, an endpoint request, or an e2e test fails, diagnose the failure, fix the implementation, and repeat the relevant checks. Do not finish while a required check is failing.
- Stop all application processes and temporary services that you started after verification, including when a check fails.

2 Vocabulary Data
- A vocabulary test consists of exactly 5 words. Each is an English word the user must translate into German.
- Use exactly these 5 English -> German pairs, stored verbatim (capitalization and umlauts matter):
  squirrel      -> Eichhörnchen
  turtle        -> Schildkröte
  refrigerator  -> Kühlschrank
  toothbrush    -> Zahnbürste
  Germany       -> Deutschland
- The same 5 words are used for every destination that offers the test. Seed this data (e.g. a Flyway migration) for the German-speaking destinations Lübeck and Vienna.

3 Discovery on the Holiday Card
- A destination that has vocabulary data shows a vocabulary icon/link on its holiday card in the holidays overview, exactly like the quiz icon. Drive this off the presence of vocabulary data — do NOT hardcode a list of cities.
- The control is an accessible link with the accessible name "Vocabulary Test".
- Destinations without vocabulary data must not show it.

4 The Vocabulary Test
- Opening the "Vocabulary Test" link shows the 5 questions and a status area. Navigation, dialog, or inline panel — your choice.
- For each question, show the English word and provide a free-text input for the German translation. Each input's accessible name must contain its English word (e.g. an input labelled "squirrel"), so every input is individually identifiable regardless of order or layout.
- A status area shows three counts, each exposed with exactly these accessible names (value included): "Unanswered: N", "Correct: N", "Incorrect: N".
- The three counts are mutually exclusive and always sum to 5. Initial state: Unanswered: 5, Correct: 0, Incorrect: 0.

5 Answering
- An answer is evaluated when the user leaves (blurs) a non-empty input. There is no submit button — the counts update live on blur.
- Blurring an empty input does nothing; the question stays unanswered.
- Comparison is exact and case- and umlaut-sensitive: the typed text must exactly equal the stored German word. "Kühlschrank" is correct; "kühlschrank", "Kuhlschrank", and "Kühlschrank " are all incorrect. No trimming or normalization.
- Once a non-empty input has been evaluated it is locked (cannot be edited again). The counts update accordingly: Unanswered decreases by 1, and Correct or Incorrect increases by 1.
15
Soverius AI

Geschwindigkeit: Hardware

LOKALE MODELLE FÜR KI-GESTÜTZTES CODING 01 WARUM Warum lokale Modelle? 02 WIE llama.cpp + Gemma 4 03 DEMO Capstone 04 PROBLEME Geschwindigkeit und Qualität GESCHWINDIGKEIT Hardware QUALITÄT Loop / Graph Engineering 16
Soverius AI

Inferenz auf CPU und GPU

CPU1× RechenleistungIntel Core i9-14900K
Festplatte oder SSD≈0,1–7 GB/sBarraCuda HDD
870 EVO / 990 PRO SSD
Swap · Keine Berechnung
Arbeitsspeicher≈90–100 GB/sDual-Channel DDR5-5600
MöglichOft zu langsam
GPU≈30–100× RechenleistungRTX 4070 Super – RTX 5090
VRAM≈500–1.800 GB/sGDDR6X – GDDR7
PraxistauglichMeist 5–10× schneller

Ein Modell mit über 1 TB lässt sich von Festplatte betreiben. Auf die Antwort wartet man entsprechend lange.

17
Soverius AI

Hardware für lokale Inferenz

Beispielkonfigurationen, keine Kaufempfehlungen.

18
Soverius AI

Einsatzbereiche

Modellgröße nach Einsatzbereich

Werte geschlossener Modelle sind Schätzungen aus öffentlichen Benchmark-Daten [1][2]. Parameterzahlen sind nicht offengelegt. Typischer Fehler: Faktor ≈2.

19
Soverius AI

Die Gemma-4-Familie

DenseGleicher Pfad für jedes TokenMixture of ExpertsRouter wählt Experten aus
Parameter insgesamt 30B 20B 10B 0
E2B5,1B gesamtDense
2,3B effektiv
E4B8B gesamtDense
4,5B effektiv
12B11,95B gesamtDense
26B-A4B25,2B gesamtMoE
3,8B aktiv
31B30,7B gesamtDense

EKerngröße ohne große Embedding-Tabellen

APro Token vom Router ausgewählte Parameter

20
Soverius AI

Speicherbedarf

EmpfohlenVollständig im VRAM

Einflussfaktoren

  • Modellgröße
  • Quantisierung
  • KV Cache
  • Kontextgröße
  • Betriebsspeicher
21
Soverius AI

Das Prinzip der Quantisierung

22
Soverius AI

Die passende Quantisierung

23
Soverius AI

Videokompression im Zeitverlauf

Beispielhafte Dateigröße für eine Minute desselben 1080p-Videos bei vergleichbarer Bildqualität.

24
Soverius AI

Prefill und Decoding ohne KV Cache

0:00.0 / 0:18.5

Prefill

Rechenleistung limitiert

Decode

Berechnet Kontext neu
  • Nutzereingabe
  • Tool-Antwort
  • Thinking
BitteerstelleeinegetesteteKomponente
mitunserembestehendenSignalStore
undergänzedieBackendIntegration
Modellschichten
Vorhergesagtes Token
HierdasfertigeErgebnis
EingabeHierdasfertige
25
Soverius AI

KV Cache während der Inferenz

0:00.0 / 0:12.5

Prefill

Baut den Cache auf

Decode

Nutzt den Cache erneut
  • Nutzereingabe
  • Tool-Antwort
  • Thinking
KV Cache
BitteerstelleeinegetesteteKomponente
mitunserembestehendenSignalStore
undergänzedieBackendIntegration
Modellschichten
Vorhergesagtes Token
HierdasfertigeErgebnis
EingabeHierdasfertige
26
Soverius AI

Speicherbedarf: Gemma 4 26B-A4B

IT QAT UD-Q4_K_XL von unsloth · 3,8B aktiv / 25,2B gesamt · ein Inferenz-Slot · F16 KV Cache

Alle 25,2B Parameter müssen in den Speicher passen. Summen enthalten 2 GiB Runtime-Reserve.

27
Soverius AI

Speicherbedarf: Qwen 3.8 27B

UD-Q4_K_XL von unsloth · 27B Dense · 16 KV-Schichten · ein Inferenz-Slot · F16 KV Cache

Hybride Architektur: 16 von 64 Schichten vergrößern den KV Cache. Summen enthalten 2 GiB Runtime-Reserve.

28
Soverius AI

Qualität: Loop / Graph Engineering

LOKALE MODELLE FÜR KI-GESTÜTZTES CODING 01 WARUM Warum lokale Modelle? 02 WIE llama.cpp + Gemma 4 03 DEMO Capstone 04 PROBLEME Geschwindigkeit und Qualität GESCHWINDIGKEIT Hardware QUALITÄT Loop / Graph Engineering 29
Soverius AI

Das umgebende System

Engineering rund um das Modell

Modelle machen Fehler. Das umgebende System muss sie erkennen und die Kontrolle behalten.

  • Jeden Modelldurchlauf begrenzen und extern prüfen, damit interne Loops den Workflow nicht blockieren.
  • Eine gut strukturierte Codebasis gibt dem Modell verlässliche Muster vor.
  • Projektspezifische Instructions und Skills vermitteln Angular, Spring, Playwright und die Projektkonventionen.
  • Lokale Inferenz rund um die Uhr ermöglicht Loop / Graph Engineering mit unabhängigen Reviews und wiederholter Validierung.
30
Soverius AI

Unser Loop

31
Soverius AI

Ein praktischer Einstieg

32
Soverius AI

Vielen Dank

Fragen?

Lokale Modelle für KI-gestütztes Coding

Dein Feedback zum Vortrag