LIFEHUBBER
Choose theme

AI Resources

FreeToken

GitHub stars: 14K GitHub forks: 1.4K Declared license: Apache-2.0: Apache-2.0 Last pushed September 30, 2026: Pushed today
Stats from GitHub

FreeToken runs large mixture-of-experts language models on a personal computer by sharing the work across the GPU, system memory, and CPU.

It is the engine between downloaded model weights and the chat app or coding agent you want to use. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

A runtime for models larger than GPU memory

A mixture-of-experts model activates only some of its expert weights for each token. FreeToken uses that pattern to keep some weights in RAM and bring in, or compute, the parts needed next.

Why it stands out

Memory can change jobs during a session

The project describes reallocating GPU memory between cached experts and conversation context while the engine is running, without restarting it or reloading the weights.

Availability

Desktop and command-line paths

The Apache-2.0 repository links Windows and Linux desktop downloads, a Python package, and source installation. The Python path has its own hardware and CUDA requirements.

Why it matters

What makes it useful

A model that exceeds your graphics card's memory may still be a candidate for this kind of setup. FreeToken's CPU-GPU execution gives you another route to explore before replacing the GPU; the available RAM and transfer bandwidth still constrain what is practical.

Notable points

What stands out

Its context cache is designed for agent sessions where tool calls and thinking blocks change the conversation. The project describes saving checkpoints so these edits need not trigger the same context computation from the beginning.

Before using

What to review

The Python installation guide lists Linux x86_64, NVIDIA hardware, driver r580 or newer, Python 3.10 or newer, and a CUDA 13 toolkit for first-use kernel compilation.

Check the supported-model and backend guide before downloading weights. The model still needs storage and enough memory across the machine.

The coding-agent launcher can write provider configuration and install a missing CLI. Its dry-run option shows the proposed changes first.

Reader fit

Who may find it relevant

For local-AI builders comfortable with model downloads, drivers, and runtime settings. The useful starting question is whether an explicitly supported MoE model fits the whole computer's resources, rather than the GPU alone.

Editorial note

Why LifeHubber lists it

FreeToken makes the hardware side of local AI easier to reason about: model size, active work per token, and GPU capacity are different constraints. That distinction helps explain why a large MoE model can have a different serving path from a similarly sized dense model.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

What to explore next

Compare the engine with the model and hardware choices around it.

FreeToken changes how a large MoE model uses one machine. These next paths help compare a different low-memory method and the wider local setup before choosing a runtime.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving