Choose theme
AI Resources
FreeToken
FreeToken runs large mixture-of-experts language models on a personal computer by sharing the work across the GPU, system memory, and CPU.
It is the engine between downloaded model weights and the chat app or coding agent you want to use. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
A runtime for models larger than GPU memory
A mixture-of-experts model activates only some of its expert weights for each token. FreeToken uses that pattern to keep some weights in RAM and bring in, or compute, the parts needed next.
Why it stands out
Memory can change jobs during a session
The project describes reallocating GPU memory between cached experts and conversation context while the engine is running, without restarting it or reloading the weights.
Availability
Desktop and command-line paths
The Apache-2.0 repository links Windows and Linux desktop downloads, a Python package, and source installation. The Python path has its own hardware and CUDA requirements.
Why it matters
What makes it useful
A model that exceeds your graphics card's memory may still be a candidate for this kind of setup. FreeToken's CPU-GPU execution gives you another route to explore before replacing the GPU; the available RAM and transfer bandwidth still constrain what is practical.
What to know
Where it fits
Start a model server, then point an OpenAI- or Anthropic-compatible client at its local address. A terminal chat interface is included. Coding-agent launch helpers can also configure clients to use that server.
Notable points
What stands out
Its context cache is designed for agent sessions where tool calls and thinking blocks change the conversation. The project describes saving checkpoints so these edits need not trigger the same context computation from the beginning.
Before using
What to review
The Python installation guide lists Linux x86_64, NVIDIA hardware, driver r580 or newer, Python 3.10 or newer, and a CUDA 13 toolkit for first-use kernel compilation.
Check the supported-model and backend guide before downloading weights. The model still needs storage and enough memory across the machine.
The coding-agent launcher can write provider configuration and install a missing CLI. Its dry-run option shows the proposed changes first.
Reader fit
Who may find it relevant
For local-AI builders comfortable with model downloads, drivers, and runtime settings. The useful starting question is whether an explicitly supported MoE model fits the whole computer's resources, rather than the GPU alone.
Editorial note
Why LifeHubber lists it
FreeToken makes the hardware side of local AI easier to reason about: model size, active work per token, and GPU capacity are different constraints. That distinction helps explain why a large MoE model can have a different serving path from a similarly sized dense model.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
What to explore next
Compare the engine with the model and hardware choices around it.
FreeToken changes how a large MoE model uses one machine. These next paths help compare a different low-memory method and the wider local setup before choosing a runtime.
More in Ecosystem
Keep browsing this category
Explore more AI ecosystem resources.
Laya
NandhaKishorM/laya
An early Apache-2.0 model family and Python runtime for bounded choice, score, and yes-or-no decisions, with English, multilingual, and task-specialized checkpoints plus a router that selects between them.
CLM
Contrastive-LM/CLM
An Apache-2.0 contrastive model and local server for typed decisions and candidate ranking, with separate state and action embeddings, reusable candidate caches, a browser playground, public reference heads, and fine-tuning tools.
Voicebox
jamiepine/voicebox
A local-first voice synthesis studio for voice cloning, speech generation, effects, and voice-powered app workflows.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.