LLM & Models

Kimi K3 and Qwen 3.8-Max: open-weights models cross 2 trillion parameters

Open weights just changed scale. Moonshot (Kimi) and Alibaba (Qwen) have each published a model with more than 2 trillion parameters.

Kimi K3

  • 2.8T parameters, Mixture-of-Experts with 104B active per token (16 of 896 experts).
  • Natively multimodal (text + image), 1-million-token context.
  • Weights on Hugging Face, quantized in MXFP4, under a custom "Kimi K3 License" (not a standard open-source license: read it before any commercial use).
  • Per its technical report, it still trails the top proprietary models (Claude Fable 5, GPT-5.6 Sol) but beats the other models in its test suite.
  • Official API: $3 per million input tokens ($0.30 cached), $15 output.

Qwen 3.8-Max

  • 2.4T parameters, 95B active, announced August 2 as the first "Max"-class model with open weights (release promised for the following week, on Hugging Face and ModelScope).
  • Focused on coding, knowledge work and long-horizon tasks.

We haven't re-verified that Qwen 3.8-Max's weights are actually live: confirm on the official page before planning anything.

Can you self-host them?

Not on a VPS. As an order of magnitude, 2.8T parameters at 4 bits is already around 1.4 TB of weights before context cache, i.e. several datacenter GPUs. That's our own estimate, not a vendor figure.

The practical value lies elsewhere:

  • Third-party providers: hosts serve these models over API, avoiding single-vendor dependence.
  • Price pressure: a near-frontier open model drags prices down.
  • Smaller distilled models: the ecosystem typically derives locally runnable versions, worth watching.

For realistic self-hosting today, a gateway like LiteLLM in front of a mix of smaller local models and APIs remains the most sensible setup.

Sources: Kimi K3 technical report, Hugging Face model card, Kimi blog, Qwen blog.