{"ok":true,"source":"tensorfeed.ai","lastUpdated":"2026-09-14","count":39,"models":[{"id":"deepseek-v4.1-flash","name":"DeepSeek V4.1 Flash","family":"DeepSeek","activeParamsB":16,"totalParamsB":552,"contextWindow":1000000,"released":"2026-09","license":"MIT","hfUrl":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash","url":"https://api-docs.deepseek.com/updates","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"First model on DeepSeek's Causal Encoder-Decoder architecture and, despite the Flash name, a far bigger download than V4 Flash: 510GB of weights against 160GB. 552B backbone plus a 196B Engram conditional memory table, activating 8B parameters per token in prefill and 16B in decode. Native image input, 1M context, and a global KV cache of about 890 bytes per token, roughly a quarter of V4 Flash, which is the case for running it on input-heavy agent loops. Plain MIT. The DeepSeek API retired V4 Flash in its favor on 2026-09-10, and DeepSeek reports Terminal-Bench 2.1 at 90.6.","quantizations":[{"id":"native","name":"FP4/FP8 (native)","vramGB":515,"quality":100,"recommendedGpu":"4x B200","notes":"DeepSeek reference release, 510GB of safetensors"}]},{"id":"minicpm5-2b","name":"MiniCPM5-2B","family":"OpenBMB","activeParamsB":2.52,"totalParamsB":2.52,"contextWindow":131072,"released":"2026-09","license":"Apache-2.0","hfUrl":"https://huggingface.co/openbmb/MiniCPM5-2B","url":"https://huggingface.co/openbmb/MiniCPM5-2B","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"Dense 2.5B (under 2B non-embedding) with 131K native context, built for local assistants and small coding and tool-use agents. Plain Apache 2.0, and OpenBMB published the pretraining, SFT, agent, and RL datasets behind it, which makes it one of the more reproducible small models. Official GGUF, GPTQ, and MLX builds ship alongside the weights, so laptop-class hardware is enough.","quantizations":[{"id":"fp16","name":"BF16","vramGB":5.1,"quality":100,"recommendedGpu":"1x 8GB consumer GPU","notes":"Reference precision"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":1.6,"quality":93,"recommendedGpu":"CPU or 4GB GPU","notes":"Official GGUF build"}]},{"id":"hy4-preview","name":"Hy4 preview","family":"Tencent","activeParamsB":49,"totalParamsB":770,"contextWindow":1000000,"released":"2026-08","license":"Apache-2.0","hfUrl":"https://huggingface.co/tencent/Hy4-preview","url":"https://hunyuan.tencent.com","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"Tencent's next flagship, shipped early under plain Apache 2.0. 770B total with 49B active across 256 routed experts, Gated DeepSeek Sparse Attention, 1M context, and a built-in MTP layer for speculative decoding. Tencent itself lists over-long reasoning and a habit of over-verifying its own work as known issues, so treat it as an evaluation target rather than a production default until the final release.","quantizations":[{"id":"fp16","name":"BF16","vramGB":1560,"quality":100,"recommendedGpu":"16x B200 (multi-node)","notes":"Official BF16 checkpoint"},{"id":"fp8","name":"FP8","vramGB":815,"quality":99,"recommendedGpu":"8x H200","notes":"Official FP8 checkpoint"}]},{"id":"granite-4.2-30b","name":"Granite 4.2 30B","family":"IBM","activeParamsB":29.3,"totalParamsB":29.3,"contextWindow":128000,"released":"2026-08","license":"Apache-2.0","hfUrl":"https://huggingface.co/ibm-granite/granite-4.2-30b","url":"https://github.com/ibm-granite/granite-4.2-language-models","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"IBM's first Granite generation with native thinking, released 2026-08-25 in 3B, 8B, and 30B dense sizes under Apache 2.0. The 30B is post-trained from Granite 4.1 30B Base with an extra agentic SFT pass. 128K native context, extendable to 512K. IBM ships official FP8, NVFP4, GGUF, and MLX builds, so it is the low-friction enterprise pick when license and provenance paperwork matter more than leaderboard rank.","quantizations":[{"id":"fp16","name":"BF16","vramGB":60,"quality":100,"recommendedGpu":"1x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":31,"quality":99,"recommendedGpu":"1x RTX 6000 Ada","notes":"Official FP8 checkpoint"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":17,"quality":93,"recommendedGpu":"1x RTX 4090","notes":"Official GGUF build"}]},{"id":"glm-5.3-flash","name":"GLM-5.3-Flash","family":"Zhipu","activeParamsB":18,"totalParamsB":320,"contextWindow":1000000,"released":"2026-08","license":"MIT","hfUrl":"https://huggingface.co/zai-org/GLM-5.3-Flash","url":"https://z.ai","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"First natively multimodal model in the GLM-5 series and the cheapest plain-MIT path to frontier agentic coding. 320B total with 18B active on a hybrid sparse plus linear attention MoE. Published 2026-08-26 after a week of anonymous preview as Ox Alpha. Terminal-Bench 2.1 84.3 and DeepSWE 63.4 put it within a few points of the unreleased flagship at roughly a tenth of the price.","quantizations":[{"id":"fp16","name":"FP16","vramGB":660,"quality":100,"recommendedGpu":"8x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":340,"quality":99,"recommendedGpu":"4x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":170,"quality":95,"recommendedGpu":"2x H200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":180,"quality":93,"recommendedGpu":"2x H200","notes":"llama.cpp / Ollama path"}]},{"id":"glm-5.3","name":"GLM-5.3","family":"Zhipu","activeParamsB":40,"totalParamsB":753,"contextWindow":1000000,"released":"2026-08","license":"GLM-5.3 License","hfUrl":"https://huggingface.co/zai-org/GLM-5.3","url":"https://z.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"Weights are now public: FP8 in the main repo and a separate BF16 repo, 753B parameters on the same base as GLM-5.2. Every gain came from post-training, which is why Terminal-Bench 3.0 jumped from 4.6 to 28.3. Check the license before deploying: unlike the plain-MIT GLM-5.2 and GLM-5.3-Flash, the GLM-5.3 License is MIT-style but requires any Model-as-a-Service operator with over $10B in trailing 12-month revenue to pass a Z.ai security review before commercial use. Z.ai also flags cyber capability as an emergent strength (CyberGym 84.5).","quantizations":[{"id":"fp16","name":"BF16","vramGB":1510,"quality":100,"recommendedGpu":"16x B200 (multi-node)","notes":"Official BF16 repo"},{"id":"fp8","name":"FP8","vramGB":760,"quality":99,"recommendedGpu":"8x H200","notes":"Official release precision"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":430,"quality":93,"recommendedGpu":"4x H200","notes":"Community GGUF (unsloth)"}]},{"id":"qwen3.8-2.4t-a95b","name":"Qwen3.8 2.4T-A95B","family":"Qwen","activeParamsB":95,"totalParamsB":2400,"contextWindow":262144,"released":"2026-08","license":"Qwen3.8-Max License","hfUrl":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B","url":"https://qwen.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"The largest open-weight model published to date and the open counterpart to the hosted Qwen3.8-Max. 2.4T total across 512 experts with 95B active. Text only, with thinking mode permanently on. Note the license is a custom Qwen3.8-Max License, not the Apache 2.0 that covers the 27B. Context is 262,144 native and stretches to roughly 1M with YaRN. This needs a multi-node rack, not a server.","quantizations":[{"id":"fp16","name":"FP16","vramGB":4945,"quality":100,"recommendedGpu":"16x B200 (multi-node)","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":2545,"quality":99,"recommendedGpu":"16x B200 (multi-node)","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":1270,"quality":95,"recommendedGpu":"8x B200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":1355,"quality":93,"recommendedGpu":"8x B200","notes":"llama.cpp / Ollama path"}]},{"id":"qwen3.8-27b","name":"Qwen3.8 27B","family":"Qwen","activeParamsB":27,"totalParamsB":27,"contextWindow":262144,"released":"2026-08","license":"Apache-2.0","hfUrl":"https://huggingface.co/Qwen/Qwen3.8-27B","url":"https://qwen.ai","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"The download leader of the 2026 open-weight wave at roughly 3.3M pulls a month. Dense 27B native vision-language model taking text, image, and video in, under plain Apache 2.0 with switchable thinking. Fits one H100 at FP8 and a single 4090 at Q4, which is the reason for the volume. OSWorld-Verified 84.3 and LiveCodeBench v6 90.3 are unusual for a model this size.","quantizations":[{"id":"fp16","name":"FP16","vramGB":56,"quality":100,"recommendedGpu":"1x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":29,"quality":99,"recommendedGpu":"1x RTX 6000 Ada","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":14,"quality":95,"recommendedGpu":"1x RTX 4090","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":15,"quality":93,"recommendedGpu":"1x RTX 4090","notes":"llama.cpp / Ollama path"}]},{"id":"qwen3.8-flash-next","name":"Qwen3.8-Flash-Next","family":"Qwen","activeParamsB":6,"totalParamsB":125,"contextWindow":262144,"released":"2026-08","license":"Qwen Community 1.0","hfUrl":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next","url":"https://qwen.ai","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"Multimodal 125B-A6B MoE with hybrid QSA attention and n-gram embeddings, the Flash tier of the Qwen3.8 line. Only 6B active per token, so it serves far cheaper than its total size suggests. Size VRAM off the checkpoint, not the headline: the 125B backbone ships with another 51B of n-gram embeddings and a 4B MTP module, 180B in all, though Qwen notes the embeddings offload more easily than experts do.","quantizations":[{"id":"fp16","name":"BF16","vramGB":365,"quality":100,"recommendedGpu":"4x H200","notes":"Full 180B checkpoint including n-gram embeddings"},{"id":"fp8","name":"FP8","vramGB":190,"quality":99,"recommendedGpu":"2x H200","notes":"Official Qwen FP8 checkpoint"},{"id":"nvfp4","name":"NVFP4","vramGB":135,"quality":96,"recommendedGpu":"1x B200","notes":"NVIDIA 4-bit build; full speedup needs Blackwell"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":105,"quality":93,"recommendedGpu":"2x H100-80GB","notes":"Community GGUF (unsloth)"}]},{"id":"nemotron-3.5-lightning","name":"Nemotron 3.5 Lightning 30B-A3B","family":"NVIDIA","activeParamsB":3,"totalParamsB":30,"contextWindow":1000000,"released":"2026-08","license":"OpenMDW-1.1","hfUrl":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16","url":"https://www.nvidia.com","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"Hybrid Mamba-2 plus MoE at 30B total and 3B active, shipped as a customization and post-training base rather than a finished assistant. Released under OpenMDW-1.1, an external open-model license, not an NVIDIA-authored one. The card claims up to 1M context but notes roughly 256K is practical on a single H100.","quantizations":[{"id":"fp16","name":"FP16","vramGB":62,"quality":100,"recommendedGpu":"1x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":32,"quality":99,"recommendedGpu":"1x RTX 6000 Ada","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":16,"quality":95,"recommendedGpu":"1x RTX 4090","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":17,"quality":93,"recommendedGpu":"1x RTX 4090","notes":"llama.cpp / Ollama path"}]},{"id":"muse-glimmer-30b","name":"Muse Glimmer 30B","family":"Meta","activeParamsB":30,"totalParamsB":30,"contextWindow":131072,"released":"2026-08","license":"Apache-2.0","hfUrl":"https://huggingface.co/meta-models/Muse-Glimmer-30B","url":"https://research.meta.ai","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"Meta Superintelligence Labs' dense multimodal 30B built for always-on local agents on a single consumer GPU, under plain Apache 2.0. SWE-Bench Verified 76.0 and AIME 2026 94.7 at high reasoning are strong for a dense model at this size. Note the repo lives under the meta-models org, not facebook or meta-llama.","quantizations":[{"id":"fp16","name":"FP16","vramGB":62,"quality":100,"recommendedGpu":"1x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":32,"quality":99,"recommendedGpu":"1x RTX 6000 Ada","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":16,"quality":95,"recommendedGpu":"1x RTX 4090","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":17,"quality":93,"recommendedGpu":"1x RTX 4090","notes":"llama.cpp / Ollama path"}]},{"id":"ling-3.0-flash","name":"Ling-3.0-flash","family":"InclusionAI","activeParamsB":5.1,"totalParamsB":124,"contextWindow":262144,"released":"2026-07","license":"MIT","hfUrl":"https://huggingface.co/inclusionAI/Ling-3.0-flash","url":"https://www.inclusionai.org","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"Ant Group's 124B hybrid-reasoning MoE with only 5.1B active, alternating Kimi Delta Attention and MLA for production-scale agentic inference. Plain MIT. The 1/64 sparsity is the point: it serves at roughly the cost of a 5B model. FP8, FP4, and INT4 checkpoints ship alongside the reference weights.","quantizations":[{"id":"fp16","name":"FP16","vramGB":255,"quality":100,"recommendedGpu":"2x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":130,"quality":99,"recommendedGpu":"1x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":66,"quality":95,"recommendedGpu":"1x H100-80GB","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":70,"quality":93,"recommendedGpu":"1x H100-80GB","notes":"llama.cpp / Ollama path"}]},{"id":"hy3","name":"Hy3","family":"Tencent","activeParamsB":21,"totalParamsB":295,"contextWindow":262144,"released":"2026-07","license":"Apache-2.0","hfUrl":"https://huggingface.co/tencent/Hy3","url":"https://hunyuan.tencent.com","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"295B MoE with 21B active plus a 3.8B multi-token-prediction layer. Worth flagging for license reasons: the April preview shipped under a restrictive license carving out the EU, UK, and South Korea, and the final release replaced that with plain Apache 2.0. SWE-bench Verified 78.0 and GPQA Diamond 90.4. Distinct from the Hy-MT2 translation line.","quantizations":[{"id":"fp16","name":"FP16","vramGB":610,"quality":100,"recommendedGpu":"8x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":315,"quality":99,"recommendedGpu":"4x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":155,"quality":95,"recommendedGpu":"2x H200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":165,"quality":93,"recommendedGpu":"2x H200","notes":"llama.cpp / Ollama path"}]},{"id":"laguna-s-2.1","name":"Laguna S 2.1","family":"poolside","activeParamsB":8,"totalParamsB":118,"contextWindow":1048576,"released":"2026-07","license":"OpenMDW-1.1","hfUrl":"https://huggingface.co/poolside/Laguna-S-2.1","url":"https://poolside.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"118B total with roughly 8B active and a full 1M context, positioned as the West's most capable open-weight coding model. Terminal-Bench 2.1 70.2 and SWE-bench Multilingual 78.5 beat several models many times its size. OpenMDW-1.1 licensed.","quantizations":[{"id":"fp16","name":"FP16","vramGB":245,"quality":100,"recommendedGpu":"2x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":125,"quality":99,"recommendedGpu":"1x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":63,"quality":95,"recommendedGpu":"1x H100-80GB","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":67,"quality":93,"recommendedGpu":"1x H100-80GB","notes":"llama.cpp / Ollama path"}]},{"id":"solar-open2-250b","name":"Solar Open 2 250B","family":"Upstage","activeParamsB":15,"totalParamsB":250,"contextWindow":1000000,"released":"2026-07","license":"Upstage Solar License","hfUrl":"https://huggingface.co/upstage/Solar-Open2-250B","url":"https://www.upstage.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"250B hybrid-attention MoE with 15B active across 321 experts, trained on roughly 12T tokens for agentic and document-heavy enterprise work. The strongest Korean-language coverage in this catalog (KMMLU-Pro 78.4, CLIcK 90.7) alongside SWE-Bench Verified 70.4. The Solar License is an Apache 2.0 derivative permitting commercial use.","quantizations":[{"id":"fp16","name":"FP16","vramGB":515,"quality":100,"recommendedGpu":"4x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":265,"quality":99,"recommendedGpu":"2x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":130,"quality":95,"recommendedGpu":"1x H200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":140,"quality":93,"recommendedGpu":"1x H200","notes":"llama.cpp / Ollama path"}]},{"id":"ornith-1.5-397b","name":"Ornith-1.5 397B","family":"Ornith","activeParamsB":null,"totalParamsB":397,"contextWindow":262144,"released":"2026-08","license":"MIT","hfUrl":"https://huggingface.co/ornith-ai/Ornith-1.5-397B","url":"https://huggingface.co/ornith-ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"The loudest new entrant of August 2026, taking several of the top Hugging Face trending slots within days. 397B MoE under plain MIT, trained with a self-improvement loop, with 35B-A3B and 9B siblings covering the smaller tiers. Active parameter count is not published. Treat the vendor's SWE-bench Verified 86 as unverified until an independent run lands, and note the team behind Ornith is not disclosed on the card.","quantizations":[{"id":"fp16","name":"FP16","vramGB":820,"quality":100,"recommendedGpu":"8x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":420,"quality":99,"recommendedGpu":"4x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":210,"quality":95,"recommendedGpu":"2x H200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":225,"quality":93,"recommendedGpu":"2x H200","notes":"llama.cpp / Ollama path"}]},{"id":"motif-3","name":"Motif 3","family":"Motif Technologies","activeParamsB":13.2,"totalParamsB":314,"contextWindow":262144,"released":"2026-08","license":"MIT","hfUrl":"https://huggingface.co/Motif-Technologies/Motif-3","url":"https://www.motiftech.io","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"South Korea's sovereign-AI entry at frontier scale: roughly 314B total with 13.2B active, upgraded from a non-commercial preview license to plain MIT on release. SWE-bench Verified 76.2 and Terminal-Bench 2.1 74.9.","quantizations":[{"id":"fp16","name":"FP16","vramGB":645,"quality":100,"recommendedGpu":"8x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":335,"quality":99,"recommendedGpu":"4x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":165,"quality":95,"recommendedGpu":"2x H200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":175,"quality":93,"recommendedGpu":"2x H200","notes":"llama.cpp / Ollama path"}]},{"id":"dots3-note-prev","name":"dots3-note Preview","family":"Xiaohongshu","activeParamsB":16,"totalParamsB":280,"contextWindow":524288,"released":"2026-08","license":"Apache-2.0","hfUrl":"https://huggingface.co/dots-studio/dots3-note-prev","url":"https://huggingface.co/dots-studio","capabilities":["text","vision","audio","tool-use","function-calling"],"weightsAvailable":true,"notes":"Rare combination of Apache 2.0, omni-modal input (text, image, video, and audio), and a 512K window. 280B total with 16B active, trained with TEMPO reinforcement learning. Still labelled preview, so expect the card and the checkpoint to move.","quantizations":[{"id":"fp16","name":"FP16","vramGB":575,"quality":100,"recommendedGpu":"8x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":295,"quality":99,"recommendedGpu":"4x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":150,"quality":95,"recommendedGpu":"2x H200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":160,"quality":93,"recommendedGpu":"2x H200","notes":"llama.cpp / Ollama path"}]},{"id":"lfm2.5-2.6b","name":"LFM2.5-2.6B","family":"LiquidAI","activeParamsB":2.69,"totalParamsB":2.69,"contextWindow":131072,"released":"2026-08","license":"LFM Open License v1.0","hfUrl":"https://huggingface.co/LiquidAI/LFM2.5-2.6B","url":"https://www.liquid.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"The on-device pick in this catalog: 2.69B dense hybrid convolution plus attention, trained on 34T tokens, running in under 2.5 GB. Post-trained specifically for agentic tool use, and the tool-calling scores (ToolSandbox 77.83, BFCLv4 56.88) hold up against far larger models. The LFM Open License is free below $10M annual revenue, so read it before commercial deployment.","quantizations":[{"id":"fp16","name":"FP16","vramGB":5.5,"quality":100,"recommendedGpu":"1x RTX 4090","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":2.9,"quality":99,"recommendedGpu":"1x RTX 4090","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":1.4,"quality":95,"recommendedGpu":"1x RTX 4090","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":1.5,"quality":93,"recommendedGpu":"1x RTX 4090","notes":"llama.cpp / Ollama path"}]},{"id":"hy-mt2-30b-a3b","name":"Hy-MT2-30B-A3B","family":"Tencent","activeParamsB":3,"totalParamsB":30,"contextWindow":262144,"released":"2026-05","license":"Apache-2.0","hfUrl":"https://huggingface.co/tencent/Hy-MT2-30B-A3B","url":"https://hunyuan.tencent.com","capabilities":["text","tool-use"],"weightsAvailable":true,"notes":"Specialist rather than generalist: an Apache 2.0 machine translation MoE covering 33 languages that follows natural-language translation instructions. 30B total, 3B active, with 1.8B and 7B siblings. Useful when a translation step should not burn a frontier model's tokens.","quantizations":[{"id":"fp16","name":"FP16","vramGB":62,"quality":100,"recommendedGpu":"1x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":32,"quality":99,"recommendedGpu":"1x RTX 6000 Ada","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":16,"quality":95,"recommendedGpu":"1x RTX 4090","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":17,"quality":93,"recommendedGpu":"1x RTX 4090","notes":"llama.cpp / Ollama path"}]},{"id":"leanstral-1.5","name":"Leanstral 1.5 119B-A6B","family":"Mistral","activeParamsB":6.5,"totalParamsB":119,"contextWindow":262144,"released":"2026-07","license":"Apache-2.0","hfUrl":"https://huggingface.co/mistralai/Leanstral-1.5-119B-A6B","url":"https://mistral.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"Narrow and deep: a Lean 4 theorem-proving and formal-verification agent, not a general assistant. Solves 587 of 672 PutnamBench problems and saturates miniF2F. Low download volume because the audience is small, but nothing else in this catalog does formal proof work.","quantizations":[{"id":"fp16","name":"FP16","vramGB":245,"quality":100,"recommendedGpu":"2x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":125,"quality":99,"recommendedGpu":"1x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":63,"quality":95,"recommendedGpu":"1x H100-80GB","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":67,"quality":93,"recommendedGpu":"1x H100-80GB","notes":"llama.cpp / Ollama path"}]},{"id":"kimi-k3","name":"Kimi K3","family":"Moonshot","activeParamsB":104,"totalParamsB":2800,"contextWindow":1048576,"released":"2026-07","license":"Kimi K3 License","hfUrl":"https://huggingface.co/moonshotai/Kimi-K3","url":"https://moonshot.ai","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"Largest open-weight model ever shipped. 2.8T total with 104B activated across 896 experts, 16 routed plus 2 shared per token, so roughly 3.7 percent of the network fires on any given token. Ranks third on the Artificial Analysis Intelligence Index behind Fable 5 and GPT-5.6 Sol. Announced as an API on July 16; the weights landed July 27 and have since passed 2.9M monthly downloads. Ships natively in MXFP4, so there is no FP16 checkpoint to download. Downloadable does not mean runnable here: this needs a rack, not a workstation.","quantizations":[{"id":"mxfp4","name":"MXFP4 (native)","vramGB":1450,"quality":100,"recommendedGpu":"8x B200","notes":"Moonshot ships this format; it is the reference, not a downgrade"},{"id":"gguf-q3","name":"GGUF Q3_K_M","vramGB":1120,"quality":88,"recommendedGpu":"8x H200","notes":"Community requant below native precision"},{"id":"gguf-q2","name":"GGUF Q2_K","vramGB":760,"quality":79,"recommendedGpu":"8x H100-80GB","notes":"Heavy quality loss; experimental only"}]},{"id":"glm-5.2","name":"GLM-5.2","family":"Zhipu","activeParamsB":40,"totalParamsB":753,"contextWindow":1000000,"released":"2026-06","license":"MIT","hfUrl":"https://huggingface.co/zai-org/GLM-5.2","url":"https://z.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"The strongest open-weights coding model by benchmark: 62.1 on SWE-bench Pro, above GPT-5.5 at 58.6. Plain MIT with no field-of-use carve-outs, which is rarer than the open-weights label suggests. IndexShare sparse attention keeps 1M-context inference affordable. The best license-to-capability ratio in this catalog.","quantizations":[{"id":"fp16","name":"FP16","vramGB":1550,"quality":100,"recommendedGpu":"8x B200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":800,"quality":99,"recommendedGpu":"8x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":400,"quality":95,"recommendedGpu":"4x H200","notes":"Cheapest serious self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":425,"quality":93,"recommendedGpu":"4x H200","notes":"llama.cpp path"}]},{"id":"longcat-2.0","name":"LongCat-2.0","family":"Meituan","activeParamsB":45,"totalParamsB":1600,"contextWindow":1000000,"released":"2026-06","license":"MIT","hfUrl":"https://huggingface.co/meituan-longcat/LongCat-2.0","url":"https://longcat.ai","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"MIT-licensed 1.6T MoE with dynamic activation of 33 to 56B per token, purpose-built for agentic coding. The first trillion-parameter model trained and served entirely on a 50,000-card domestic Chinese cluster. Meituan self-reports 59.5 on SWE-Bench Pro; independent verification is still pending, so treat the number as vendor-supplied.","quantizations":[{"id":"fp16","name":"FP16","vramGB":3300,"quality":100,"recommendedGpu":"16x B200 (multi-node)","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":1690,"quality":99,"recommendedGpu":"8x B200","notes":"Production self-host minimum"},{"id":"awq","name":"AWQ INT4","vramGB":870,"quality":95,"recommendedGpu":"8x H200","notes":"Quantized self-host"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":940,"quality":93,"recommendedGpu":"8x H200","notes":"llama.cpp; uncommon at this size"}]},{"id":"minimax-m3","name":"MiniMax M3","family":"MiniMax","activeParamsB":23,"totalParamsB":428,"contextWindow":1000000,"released":"2026-06","license":"MiniMax Community (commercial restrictions)","hfUrl":"https://huggingface.co/MiniMaxAI/MiniMax-M3","url":"https://www.minimax.io","capabilities":["text","vision","video","tool-use","function-calling"],"weightsAvailable":true,"notes":"Sparse-attention MoE with native image and video input. Loosened from the M2.7 terms, which banned commercial use outright without written permission, but this is still a custom community license and not Apache or MIT. Read it before you build a business on it. The most capable open multimodal option at a size a small cluster can actually hold.","quantizations":[{"id":"fp16","name":"FP16","vramGB":880,"quality":100,"recommendedGpu":"8x H200","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":450,"quality":99,"recommendedGpu":"4x H200","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":230,"quality":95,"recommendedGpu":"2x H200","notes":"Two-GPU fit"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":245,"quality":93,"recommendedGpu":"2x H200","notes":"llama.cpp path"}]},{"id":"nemotron-3-ultra","name":"Nemotron 3 Ultra 550B-A55B","family":"NVIDIA","activeParamsB":55,"totalParamsB":550,"contextWindow":262144,"released":"2026-06","license":"OpenMDW-1.1","hfUrl":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16","url":"https://developer.nvidia.com/nemotron","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"NVIDIA's own flagship open model and the largest American open-weight release. Nemotron-H hybrid Mamba and Transformer architecture with 512 routed experts activating 22 per token, 256K context. NVIDIA publishes first-party BF16 and NVFP4 checkpoints, and the NVFP4 path is the point: 4-bit inference designed for Blackwell tensor cores rather than a community requant. The company that convened the July 2026 open-weights coalition letter ships a frontier-class open model itself.","quantizations":[{"id":"bf16","name":"BF16","vramGB":1150,"quality":100,"recommendedGpu":"8x B200","notes":"Official NVIDIA checkpoint"},{"id":"fp8","name":"FP8","vramGB":590,"quality":99,"recommendedGpu":"8x H200","notes":"Production default"},{"id":"nvfp4","name":"NVFP4 (native 4-bit)","vramGB":295,"quality":96,"recommendedGpu":"4x H200","notes":"Official NVIDIA 4-bit; full speedup needs Blackwell"}]},{"id":"nemotron-3-super","name":"Nemotron 3 Super 120B-A12B","family":"NVIDIA","activeParamsB":12,"totalParamsB":120,"contextWindow":262144,"released":"2026-03","license":"NVIDIA Nemotron Open Model License","hfUrl":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16","url":"https://developer.nvidia.com/nemotron","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"The most downloaded model in the Nemotron 3 family and the sweet spot of the catalog: 120B total but only 12B active per token, so throughput tracks a small model while quality tracks a large one. 256K context. Quantized to NVFP4 it fits on a single 80GB card, which makes it the most accessible agentic model at this capability tier. Permissive enough for commercial self-hosting.","quantizations":[{"id":"bf16","name":"BF16","vramGB":250,"quality":100,"recommendedGpu":"4x H100-80GB","notes":"Official NVIDIA checkpoint"},{"id":"fp8","name":"FP8","vramGB":130,"quality":99,"recommendedGpu":"2x H100-80GB","notes":"Official NVIDIA checkpoint"},{"id":"nvfp4","name":"NVFP4 (native 4-bit)","vramGB":68,"quality":96,"recommendedGpu":"1x H100-80GB","notes":"Single-GPU production fit"}]},{"id":"nemotron-3-nano","name":"Nemotron 3 Nano 30B-A3B","family":"NVIDIA","activeParamsB":3,"totalParamsB":30,"contextWindow":262144,"released":"2025-12","license":"NVIDIA Nemotron Open Model License","hfUrl":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16","url":"https://developer.nvidia.com/nemotron","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"The workstation entry point. 30B total with 3B active across 128 experts, yet it keeps the full 256K context of its larger siblings. At NVFP4 it runs on a single consumer 24GB card, which makes it the cheapest way to put a modern long-context agent on hardware you already own. An Omni variant adds vision and audio.","quantizations":[{"id":"bf16","name":"BF16","vramGB":64,"quality":100,"recommendedGpu":"1x H100-80GB","notes":"Official NVIDIA checkpoint"},{"id":"fp8","name":"FP8","vramGB":34,"quality":99,"recommendedGpu":"1x A100-40GB","notes":"Official NVIDIA checkpoint"},{"id":"nvfp4","name":"NVFP4 (native 4-bit)","vramGB":18,"quality":96,"recommendedGpu":"1x RTX 4090","notes":"Consumer-GPU fit"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":19,"quality":93,"recommendedGpu":"1x RTX 4090","notes":"llama.cpp / Ollama"}]},{"id":"command-a-plus","name":"Command A+","family":"Cohere","activeParamsB":25,"totalParamsB":218,"contextWindow":128000,"released":"2026-05","license":"Apache-2.0","hfUrl":"https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16","url":"https://cohere.com","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"Apache-2.0 MoE from an American lab, which makes it the cleanest license-plus-jurisdiction combination in the catalog for regulated buyers. Cohere publishes official bf16, FP8, and w4a4 checkpoints, so the quantizations below are first-party rather than community requants. Strong RAG and enterprise retrieval fit.","quantizations":[{"id":"bf16","name":"BF16","vramGB":460,"quality":100,"recommendedGpu":"8x H100-80GB","notes":"Official Cohere checkpoint"},{"id":"fp8","name":"FP8","vramGB":235,"quality":99,"recommendedGpu":"4x H100-80GB","notes":"Official Cohere checkpoint"},{"id":"w4a4","name":"W4A4","vramGB":120,"quality":94,"recommendedGpu":"2x H100-80GB","notes":"Official 4-bit weight and activation quant"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":125,"quality":93,"recommendedGpu":"2x H100-80GB","notes":"Community llama.cpp build"}]},{"id":"mistral-medium-3.5","name":"Mistral Medium 3.5","family":"Mistral","activeParamsB":null,"totalParamsB":128,"contextWindow":256000,"released":"2026-05","license":"Modified MIT","hfUrl":"https://huggingface.co/mistralai/Mistral-Medium-3.5-128B","url":"https://mistral.ai/news/mistral-medium-3-5/","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"Dense 128B hitting 77.6 percent on SWE-Bench Verified, which is frontier-adjacent coding performance at a size that fits on one node. 256K context, larger than Sonnet 4.6. The most practical entry in this catalog: strong enough to matter, small enough to actually run, and licensed permissively enough to ship.","quantizations":[{"id":"fp16","name":"FP16","vramGB":270,"quality":100,"recommendedGpu":"4x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":140,"quality":99,"recommendedGpu":"2x H100-80GB","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":72,"quality":96,"recommendedGpu":"1x H100-80GB","notes":"Single-GPU production fit"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":78,"quality":94,"recommendedGpu":"1x H200","notes":"llama.cpp / Ollama"}]},{"id":"llama-4-maverick","name":"Llama 4 Maverick","family":"Meta","activeParamsB":17,"totalParamsB":400,"contextWindow":1000000,"released":"2025-04","license":"Llama 4 Community License","hfUrl":"https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct","url":"https://ai.meta.com/blog/llama-4/","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"Meta flagship MoE. 17B active / 400B total. 1M context. Vision-native. Best fit for organizations that have multi-H100 / B200 capacity.","quantizations":[{"id":"fp16","name":"FP16","vramGB":800,"quality":100,"recommendedGpu":"8x H200","notes":"Full precision, multi-node typical"},{"id":"fp8","name":"FP8","vramGB":410,"quality":99,"recommendedGpu":"4x H200","notes":"Production default; minimal quality loss"},{"id":"awq","name":"AWQ INT4","vramGB":215,"quality":96,"recommendedGpu":"2x H100-80GB","notes":"Fits on 2-GPU node"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":240,"quality":94,"recommendedGpu":"1x B200","notes":"CPU/GPU offload via llama.cpp"}]},{"id":"llama-4-scout","name":"Llama 4 Scout","family":"Meta","activeParamsB":17,"totalParamsB":109,"contextWindow":10000000,"released":"2025-04","license":"Llama 4 Community License","hfUrl":"https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct","url":"https://ai.meta.com/blog/llama-4/","capabilities":["text","vision","tool-use","function-calling"],"weightsAvailable":true,"notes":"Smaller Llama 4 sibling. 17B active / 109B total. 10M context (industry record). Vision-native. The default open-weights agent choice.","quantizations":[{"id":"fp16","name":"FP16","vramGB":220,"quality":100,"recommendedGpu":"4x H100-80GB","notes":"Multi-GPU fit"},{"id":"fp8","name":"FP8","vramGB":115,"quality":99,"recommendedGpu":"2x H100-80GB","notes":"Production default"},{"id":"awq","name":"AWQ INT4","vramGB":60,"quality":96,"recommendedGpu":"1x H100-80GB","notes":"Single-GPU production fit"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":65,"quality":94,"recommendedGpu":"1x A100-80GB","notes":"llama.cpp / Ollama"},{"id":"gguf-q3","name":"GGUF Q3_K_M","vramGB":50,"quality":89,"recommendedGpu":"1x RTX 6000","notes":"Edge-device deploys"}]},{"id":"deepseek-v4-pro","name":"DeepSeek V4 Pro","family":"DeepSeek","activeParamsB":49,"totalParamsB":1600,"contextWindow":1000000,"released":"2026-04","license":"MIT","hfUrl":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","url":"https://api-docs.deepseek.com/updates","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"MIT-licensed frontier MoE with 1.6T total and 49B active, 1M context. The April release was a preview; the official DeepSeek-V4-Pro-0813 weights landed 2026-08-13 with a DSpark speculative decoding module attached and large agentic gains (Terminal-Bench 2.1 from 72.1 to 87.9, DeepSWE from 12.8 to 62.7 on DeepSeek's numbers). Pull the 0813 repo, not the original preview. Self-hosting needs a full 8-GPU H200 or B200 node.","quantizations":[{"id":"native","name":"FP4/FP8 (native)","vramGB":895,"quality":100,"recommendedGpu":"8x H200","notes":"DeepSeek reference release format"},{"id":"nvfp4","name":"NVFP4","vramGB":945,"quality":99,"recommendedGpu":"8x B200","notes":"NVIDIA build for Blackwell kernels; no smaller than native"}]},{"id":"deepseek-v4-flash","name":"DeepSeek V4 Flash","family":"DeepSeek","activeParamsB":13,"totalParamsB":284,"contextWindow":1000000,"released":"2026-04","license":"MIT","hfUrl":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash","url":"https://www.deepseek.com","capabilities":["text","tool-use","function-calling"],"weightsAvailable":true,"notes":"Cheap DeepSeek tier and the throughput workhorse of the MIT-licensed open stack. 284B total with only 13B active per token (top-6 of 256 routed experts plus one shared), so serving cost tracks a 13B model while quality tracks something far larger. Official release is a mixed native format: FP4 for the MoE experts, FP8 for the dense layers. Note the total parameter count, not the active one, when sizing VRAM. The DeepSeek API retired V4 Flash for V4.1 Flash on 2026-09-10, but these MIT weights stay downloadable and remain the cheaper self-host at under a third of the V4.1 download.","quantizations":[{"id":"native","name":"FP4/FP8 (native)","vramGB":160,"quality":100,"recommendedGpu":"2x H200","notes":"DeepSeek reference release format"},{"id":"fp16","name":"FP16","vramGB":600,"quality":100,"recommendedGpu":"8x H100-80GB","notes":"Upcast; rarely worth it over native"},{"id":"fp8","name":"FP8","vramGB":300,"quality":99,"recommendedGpu":"4x H100-80GB","notes":"Uniform FP8 across all layers"},{"id":"awq","name":"AWQ INT4","vramGB":155,"quality":95,"recommendedGpu":"2x H200","notes":"Comparable footprint to native FP4"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":165,"quality":93,"recommendedGpu":"2x H200","notes":"llama.cpp path"}]},{"id":"qwen-2.5-72b","name":"Qwen 2.5 72B Instruct","family":"Alibaba","activeParamsB":72,"totalParamsB":72,"contextWindow":130000,"released":"2024-09","license":"Qwen License","hfUrl":"https://huggingface.co/Qwen/Qwen2.5-72B-Instruct","url":"https://qwenlm.github.io","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"Strong multilingual (29 languages). Solid coding performance. Workhorse alternative to Llama 4 Scout for organizations that prefer dense over MoE.","quantizations":[{"id":"fp16","name":"FP16","vramGB":145,"quality":100,"recommendedGpu":"2x H100-80GB","notes":"Reference precision"},{"id":"fp8","name":"FP8","vramGB":75,"quality":99,"recommendedGpu":"1x H100-80GB","notes":"Single-GPU fit"},{"id":"awq","name":"AWQ INT4","vramGB":40,"quality":96,"recommendedGpu":"1x A100-80GB","notes":"Cheaper GPU"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":45,"quality":94,"recommendedGpu":"1x RTX 6000","notes":"Workstation"}]},{"id":"mixtral-8x22b","name":"Mixtral 8x22B Instruct","family":"Mistral","activeParamsB":39,"totalParamsB":141,"contextWindow":65536,"released":"2024-04","license":"Apache-2.0","hfUrl":"https://huggingface.co/mistralai/Mixtral-8x22B-Instruct-v0.1","url":"https://mistral.ai/news/mixtral-8x22b/","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"Apache-2.0 MoE; the cleanest open license in the catalog. Older but battle-tested. Strong fit for production deployments where license clarity matters.","quantizations":[{"id":"fp16","name":"FP16","vramGB":282,"quality":100,"recommendedGpu":"4x H100-80GB","notes":"Multi-GPU"},{"id":"fp8","name":"FP8","vramGB":145,"quality":99,"recommendedGpu":"2x H100-80GB","notes":"Production"},{"id":"awq","name":"AWQ INT4","vramGB":75,"quality":95,"recommendedGpu":"1x H100-80GB","notes":"Single-GPU fit"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":80,"quality":93,"recommendedGpu":"1x A100-80GB","notes":"llama.cpp"}]},{"id":"gemma-3-27b","name":"Gemma 3 27B Instruct","family":"Google","activeParamsB":27,"totalParamsB":27,"contextWindow":128000,"released":"2025-03","license":"Gemma Terms of Use","hfUrl":"https://huggingface.co/google/gemma-3-27b-it","url":"https://blog.google/technology/developers/gemma-3/","capabilities":["text","vision","tool-use","multilingual"],"weightsAvailable":true,"notes":"Google open model with native vision. 140 languages. Light fine-tune target; strong base for domain-specific agents.","quantizations":[{"id":"fp16","name":"FP16","vramGB":54,"quality":100,"recommendedGpu":"1x H100-80GB","notes":"Single-GPU FP16"},{"id":"fp8","name":"FP8","vramGB":28,"quality":99,"recommendedGpu":"1x A100-40GB","notes":"Cheaper GPU"},{"id":"awq","name":"AWQ INT4","vramGB":16,"quality":96,"recommendedGpu":"1x RTX 4090","notes":"Consumer GPU"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":18,"quality":94,"recommendedGpu":"1x RTX 4090","notes":"llama.cpp/Ollama"}]},{"id":"llama-3.3-70b","name":"Llama 3.3 70B Instruct","family":"Meta","activeParamsB":70,"totalParamsB":70,"contextWindow":128000,"released":"2024-12","license":"Llama 3.3 Community License","hfUrl":"https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct","url":"https://ai.meta.com/blog/meta-llama-3-3/","capabilities":["text","tool-use","function-calling","multilingual"],"weightsAvailable":true,"notes":"Late-2024 dense Llama. Stronger than Llama 3.1 70B at the same parameter count. Workhorse before Llama 4 Scout arrived in April 2025.","quantizations":[{"id":"fp16","name":"FP16","vramGB":140,"quality":100,"recommendedGpu":"2x H100-80GB","notes":"Reference"},{"id":"fp8","name":"FP8","vramGB":75,"quality":99,"recommendedGpu":"1x H100-80GB","notes":"Single-GPU"},{"id":"awq","name":"AWQ INT4","vramGB":40,"quality":96,"recommendedGpu":"1x A100-80GB","notes":"Cheaper GPU"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":45,"quality":94,"recommendedGpu":"1x RTX 6000","notes":"Workstation"}]},{"id":"phi-4","name":"Phi-4","family":"Microsoft","activeParamsB":14,"totalParamsB":14,"contextWindow":16384,"released":"2024-12","license":"MIT","hfUrl":"https://huggingface.co/microsoft/phi-4","url":"https://techcommunity.microsoft.com/blog/aiplatformblog/introducing-phi-4","capabilities":["text","tool-use"],"weightsAvailable":true,"notes":"MIT-licensed 14B with strong math performance for its size. Fits in consumer-grade VRAM. Good fit for on-device or edge agents.","quantizations":[{"id":"fp16","name":"FP16","vramGB":28,"quality":100,"recommendedGpu":"1x A100-40GB","notes":"Reference"},{"id":"fp8","name":"FP8","vramGB":15,"quality":99,"recommendedGpu":"1x RTX 4090","notes":"Consumer-GPU"},{"id":"awq","name":"AWQ INT4","vramGB":8,"quality":96,"recommendedGpu":"1x RTX 3090","notes":"Older consumer GPU"},{"id":"gguf-q4","name":"GGUF Q4_K_M","vramGB":9,"quality":94,"recommendedGpu":"1x RTX 3090","notes":"Edge / Mac M-series"}]}]}