ToolCabana
Language preview: no interface translation is available for this language yet. Showing English. Use English

Private local AI chat

Chat with a local WebGPU model or a configured AI service.

Favorites are saved in this browser. Find them in My favorites.

Private local AI chat

Your input stays in this browser unless stated otherwise
Clear your input, files and result, and restore the default settings.
Load sample text to explore what this tool can do. This replaces your current input.
0 charactersClear the source text.

How to use Private local AI chat

Private local AI chat sends a single message to a small language model, Qwen2.5 0.5B Instruct, that runs in your browser through WebGPU, and returns one written answer. The model downloads on first use and may stay cached on the device. Messages are limited to 8,000 characters and answers to 1,400 output tokens, and replies can be wrong, so review them.

  1. Paste the source into the input editor, or load the built-in example.
  2. Set processing.
  3. Run private local ai chat, review the output, then use the available copy or download controls.

What this tool supports

Downloads an inference model on first run. Chat and generation require WebGPU; smaller recognition models use browser WebAssembly. AI output can omit or invent details. Review it before use.

Processing
Local browser

Limits and processing

Chat sources: 8,000 characters; generation: up to 1,400 output tokens. First model downloads can be hundreds of MB. WebGPU chat and browser memory requirements vary by device. Model output is a draft.

Example source
A useful tool makes everyday work easier. Review the results before sharing.

Frequently asked questions

Why does the local AI chat say my browser needs WebGPU?

The chat model runs on your device's graphics hardware through WebGPU. If the browser does not expose WebGPU, the tool stops with an error before downloading the model; switch to a browser and device that support it.

How big is the model download for local AI chat?

The first run downloads the model and runtime, which can be hundreds of MB, with progress shown while it loads. Later runs can reuse the copy cached by the browser unless that cache has been cleared.

Does the local AI chat remember previous messages?

No. Each run sends only the text currently in the input box, with a short instruction to answer accurately and state uncertainty. To continue a topic, include the earlier context in your new message, staying within 8,000 characters.

Where is my input processed?

Local mode downloads public model/runtime assets and keeps input in this browser. Models may remain cached on this device.

What are the input limits?

Chat sources: 8,000 characters; generation: up to 1,400 output tokens. First model downloads can be hundreds of MB. WebGPU chat and browser memory requirements vary by device. Model output is a draft.