An external tool · tokenelli.com
Tokenelli
A full chat model running on your own GPU. Summarize pages, rewrite text, and think out loud — without anything you read or type leaving your device.
Why it belongs here
Most of the imagination economy runs through someone else’s server. That is a reasonable trade until the thing you are thinking about is the thing you cannot send anywhere — an unpublished manuscript, a client’s numbers, a diagnosis, a half-formed idea you are not ready to defend.
Tokenelli takes the other route. The model runs on your hardware, so the privacy guarantee is structural rather than contractual: there is no server to log the conversation, because there is no server in the loop at all. That changes what you are willing to bring to the machine, which is the part that matters for direction.
How it works
- 01Install the sidebar
Add the extension to Chrome or Edge. About 1 MB. The model itself downloads once, on first use.
- 02Runs on your GPU
Gemma 4 executes locally via WebGPU. Inference happens entirely in your browser, and keeps working offline once cached.
- 03You stay in control
Choose exactly what text the model sees — a selection or the whole page. Per-site permissions, nothing sent to the cloud.
Tokenelli is built and operated by a third party. We link to it because it is a good instrument, not because we run it — requirements, pricing, and availability are theirs to change. Powered by Gemma 4 E2B + WebGPU.