# Privater lokaler KI-Chat

> Ermöglicht Chats mit einem lokalen WebGPU-Modell oder einem konfigurierten KI-Dienst.

[Open tool](https://www.toolcabana.com/de/local-ai-chat) · [KI und Prompt-Tools](https://www.toolcabana.com/de/category/ai-and-prompt-tools)

Tool ID: local-ai-chat. Requested language: de. Description language: de. Complete guide translation: yes.

## Overview (en)

Private local AI chat sends a single message to a small language model, Qwen2.5 0.5B Instruct, that runs in your browser through WebGPU, and returns one written answer. The model downloads on first use and may stay cached on the device. Messages are limited to 8,000 characters and answers to 1,400 output tokens, and replies can be wrong, so review them.

## Supported tasks (de)

Lädt beim ersten Durchlauf ein Inferenzmodell herunter. Chat und Generierung erfordern WebGPU; kleinere Erkennungsmodelle verwenden WebAssembly im Browser. KI-Ausgaben können Details auslassen oder erfinden. Prüfe sie vor der Verwendung.

## Steps (de)

1. Füge die Quelle in den Eingabe-Editor ein oder lade das integrierte Beispiel.
2. Lege die Verarbeitung fest.
3. Führe den privaten lokalen KI-Chat aus, prüfe die Ausgabe und verwende dann die verfügbaren Bedienelemente zum Kopieren oder Herunterladen.

## Settings

- Verarbeitung (de; key: inference; type: select): Lokal im Browser [de]

## Limitations (de)

Chat-Quellen: 8.000 Zeichen; Generierung: bis zu 1.400 Ausgabe-Tokens. Die ersten Modell-Downloads können mehrere Hundert MB groß sein. WebGPU-Chat und Browser-Speicheranforderungen variieren je nach Gerät. Modellausgaben sind ein Entwurf.

## Example input

```text
A useful tool makes everyday work easier. Review the results before sharing.
```

## Privacy and connections (de)

Der lokale Modus lädt öffentliche Modell- und Laufzeit-Assets herunter und behält die Eingabe in diesem Browser. Modelle können auf diesem Gerät zwischengespeichert bleiben.

## Questions (de)

### Why does the local AI chat say my browser needs WebGPU?

The chat model runs on your device's graphics hardware through WebGPU. If the browser does not expose WebGPU, the tool stops with an error before downloading the model; switch to a browser and device that support it.

### How big is the model download for local AI chat?

The first run downloads the model and runtime, which can be hundreds of MB, with progress shown while it loads. Later runs can reuse the copy cached by the browser unless that cache has been cleared.

### Does the local AI chat remember previous messages?

No. Each run sends only the text currently in the input box, with a short instruction to answer accurately and state uncertainty. To continue a topic, include the earlier context in your new message, staying within 8,000 characters.

### Wo wird meine Eingabe verarbeitet?

Der lokale Modus lädt öffentliche Modell- und Laufzeit-Assets herunter und behält die Eingabe in diesem Browser. Modelle können auf diesem Gerät zwischengespeichert bleiben.

### Welche Eingabelimits gelten?

Chat-Quellen: 8.000 Zeichen; Generierung: bis zu 1.400 Ausgabe-Tokens. Die ersten Modell-Downloads können mehrere Hundert MB groß sein. WebGPU-Chat und Browser-Speicheranforderungen variieren je nach Gerät. Modellausgaben sind ein Entwurf.

## Related tools

- [Chat with document](https://www.toolcabana.com/de/chat-document)
- [AI flashcard maker](https://www.toolcabana.com/de/flashcard-maker)
- [AI quiz maker](https://www.toolcabana.com/de/quiz-maker)
- [AI study guide](https://www.toolcabana.com/de/study-guide)
