DeepSeek · local
DeepSeek V4 Pro 0813
DeepSeek’s official DeepSeek-V4-Pro-0813 release supersedes the preview and adds the DSpark speculative-decoding module to the V4-Pro structure. It retains the 1M-token context and exposes low, high, and max reasoning effort. The official local-serving example uses a node with four GB300 GPUs, while DeepSeek does not publish a complete workstation capacity profile; keep it out of consumer hardware recommendations.
Local setup review pending
LG AI Research · local
EXAONE 4.0 32B
LG AI Research's 32B hybrid reasoning model for Korean, English, and Spanish, with tool calling and both reasoning and non-reasoning modes. Official Transformers and local llama.cpp paths exist, but its license restricts competing-model development and complete hardware requirements are not published.
Local setup review pending
Z.ai · local
GLM-5.3 Flash
Z.ai's MIT-licensed native multimodal GLM-5 model with 320B total and 18B active parameters. Its hybrid sparse-and-linear attention architecture targets lower long-context serving cost, but the publisher does not provide a complete workstation requirement profile.
Local setup review pending
Z.ai · local
GLM-5.3
Z.ai's 753B open-weight, custom-licensed coding and long-horizon model. The publisher documents local vLLM and SGLang serving, but does not publish one complete workstation capacity profile; evaluate it as server infrastructure rather than a consumer-GPU recommendation.
Local setup review pending
OpenAI · local
gpt-oss-120b
OpenAI's 117B-total, 5.1B-active open-weight reasoning model with a 128K context window. OpenAI states that the MXFP4 weights fit on one 80GB H100-class GPU; this is a high-capacity workstation or server path, not a general consumer-GPU recommendation.
Local setup review pending
Kakao · local
Kanana-2 30B-A3B
Kakao's 30B-total, 3B-active MoE model family for agentic work, Korean-forward multilingual use, tool calling, and reasoning. Kakao provides vLLM and SGLang paths for the Instruct and Thinking variants; the native context is 32K and the model is released under the custom Kanana License.
Local setup review pending
Moonshot AI · local
Kimi K3
Moonshot AI's 2.8T-parameter native multimodal agentic model with a 1M-token context window. It has official vLLM, SGLang, and TokenSpeed serving paths, but uses a custom Kimi K3 License and has no complete first-party workstation requirement profile; assess licensing and capacity before deployment.
Local setup review pending
Meta · local
Llama 4 Maverick
Meta's 400B-total, 17B-active native multimodal MoE model for image and text understanding, coding, and general assistant work. Meta documents a 1M context window and a single H100 DGX-host deployment path; treat it as server-class infrastructure and review the custom Llama 4 Community License before a commercial rollout.
Local setup review pending
MiniMax · local
MiniMax M2.7
MiniMax's 229B open-weight, custom-licensed model for coding and agentic workflows. MiniMax publishes vLLM and SGLang serving paths; an independent 4-bit MLX artifact is available for Apple Silicon, but its target-Mac capacity and runtime behavior require local validation.
Local setup review pending
OpenAI · cloud
GPT-5
A cloud-only GPT entry for existing workloads, with no local hardware recommendation profile. OpenAI’s current model catalog labels GPT-5 a previous model and recommends the latest GPT-6 Astra for new coding, reasoning, and agentic use. Do not present GPT-5 as the default choice for a new deployment.
Cloud only
Qwen · both
Qwen-Image
Qwen-Image is the original 20B Apache 2.0 image foundation model for text-to-image generation and image editing. Its official repository now presents Qwen-Image-2.0 as the newer generation, so choose its newer weights when the current model rather than archival compatibility is the requirement. Qwen still does not publish a complete, first-party desktop requirement profile for a buying recommendation.
Local setup review pending
Qwen · local
Qwen3-235B-A22B
Qwen's Apache 2.0 Mixture-of-Experts flagship with 235B total and 22B active parameters. It natively supports 32,768 tokens; reaching 131,072 tokens requires YaRN configuration and Qwen documents local and production paths, but does not publish one complete workstation requirement profile; keep hardware selection outside the recommender until that evidence is complete.
Local setup review pending
Qwen · local
Qwen3.8 2.4T-A95B
Qwen's 2.4T-total, 95B-active text-only reasoning model with official vLLM and SGLang paths. Its Qwen3.8-Max license and mandatory thinking mode are deployment gates; the publisher's cloud Max variant adds separate features such as vision input and a default 1M context window.
Local setup review pending
OrcaRouter · local
Qwen3.8 27B Uncensored MLX
A community-derived, refusal-removed Qwen3.8 27B MLX build. Its current model card offers 2-bit, 4-bit, 6-bit, and 8-bit variants, with 4-bit as the default. It is not an official Qwen release and has no meaningful built-in guardrails; its publisher limits it to legitimate, controlled safety research and evaluation. Keep it out of consumer recommendations and production deployment unless an owner has approved a documented moderation and abuse-prevention plan.
Local setup review pending
Qwen · local
Qwen3.8 27B
Qwen's Apache 2.0, 27B native vision-language model with a 262K-token native context window. It is a distinct official base model; Apple Silicon ports and uncensored derivatives are community releases that must be evaluated separately.
Local setup review pending
Qwen · local
Qwen3.8 Flash-Next
Qwen's native multimodal MoE model with 125B main parameters, 6B active parameters, and 51B additional n-gram embedding parameters. Qwen documents vLLM and SGLang serving; its Apple Silicon MLX route needs target-machine validation because the publisher does not provide one complete desktop capacity profile.
Local setup review pending
Qwen · both
Qwen3 8B
An open-weight language model entry with separate inference, LoRA, full-fine-tuning, and prototype-watch profiles. Minimum memory values are compatibility gates, not speed estimates.
Upstage · local
Solar Open 2
Upstage's 250B-A15B open-weight MoE language model for long-horizon agentic work in Korean, English, and Japanese. Upstage publishes local Transformers and vLLM paths, but its supplied vLLM quickstart assumes eight H200/B200-class GPUs with at least 141GB each; this is server-class infrastructure, not a consumer-GPU hardware recommendation.
Local setup review pending
Wan · local
Wan2.1 T2V 1.3B
An Apache 2.0 text-to-video model for 480P generation. Wan’s official repository records 8.19GB VRAM for the 1.3B model, but does not publish a complete desktop RAM, storage, and operating-system requirement profile; hardware recommendations remain withheld.
Local setup review pending