{"ok":true,"source":"tensorfeed.ai","lastUpdated":"2026-09-14","count":14,"runs":[{"id":"soofi-s-30b-a3b","model":"Soofi S 30B-A3B","publisher":"Soofi consortium","released":"2026-07","activeParamsB":3,"totalParamsB":30,"trainingTokens":"26.68T","hardware":"NVIDIA B200 (Deutsche Telekom Industrial AI Cloud, Munich)","hardwareCount":512,"computeHours":"~253,000 B200 GPU-hours","estimatedCostMillionUSD":null,"costSource":"estimated","duration":"~7 weeks (March 24 to May 13, 2026)","openWeights":false,"url":"https://arxiv.org/abs/2607.09424","notes":"German-English hybrid Mamba-Transformer MoE built on the Nemotron 3 Nano reference architecture by a German research consortium (Fraunhofer IAIS, DFKI and others). Unusually complete compute disclosure: up to 512 B200s, about 253,000 GPU-hours and the exact training window. Weights are a gated closed-beta preview pending a final license; no dollar cost given."},{"id":"longcat-2.0","model":"LongCat-2.0","publisher":"Meituan","released":"2026-07","activeParamsB":48,"totalParamsB":1600,"trainingTokens":"35T+","hardware":"AI ASIC superpods (vendor not named)","hardwareCount":50000,"computeHours":"millions of accelerator-days (per Meituan)","estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":true,"url":"https://huggingface.co/meituan-longcat/LongCat-2.0","notes":"MIT-licensed MoE with about 48B active parameters. Meituan says pretraining ran on more than 50,000 AI ASICs (hardwareCount is that lower bound) with less memory per device than an H800, over 35T+ tokens with no rollbacks, plus hundreds of billions of 1M-context tokens. The accelerator vendor and dollar cost are not disclosed."},{"id":"nemotron-3-ultra","model":"Nemotron 3 Ultra","publisher":"NVIDIA","released":"2026-06","activeParamsB":55,"totalParamsB":550,"trainingTokens":"20T","hardware":"NVIDIA GB200 (long-context phase); NVFP4 pretraining","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":true,"url":"https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/","notes":"Hybrid Mamba-Attention MoE pretrained in NVFP4 on 20T text tokens (15T diversity phase, then 5T high-quality phase), then extended to 1M context on GB200 GPUs. NVIDIA calls it the largest demonstration of stable NVFP4 training to date. Base, post-trained and quantized checkpoints plus most training data are on Hugging Face; GPU count and GPU-hours are not disclosed."},{"id":"gpt-5.5","model":"GPT-5.5","publisher":"OpenAI","released":"2026-04","activeParamsB":null,"totalParamsB":null,"trainingTokens":"undisclosed","hardware":"NVIDIA GB200 and GB300 NVL72","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":false,"url":"https://openai.com/index/introducing-gpt-5-5/","notes":"Released April 23, 2026. OpenAI says GPT-5.5 was co-designed for, trained with and served on NVIDIA GB200 and GB300 NVL72 systems. The system card gives no parameter count, token count, cluster size or cost, so no figures are listed here."},{"id":"deepseek-v4-pro","model":"DeepSeek V4 Pro","publisher":"DeepSeek","released":"2026-04","activeParamsB":49,"totalParamsB":1600,"trainingTokens":"33T","hardware":"undisclosed (report validates kernels on NVIDIA GPUs and Huawei Ascend NPUs)","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":true,"url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro","notes":"MIT-licensed preview with a 1M-token context. Sibling DeepSeek V4 Flash is 284B total / 13B active on 32T tokens. The technical report details the architecture (hybrid CSA/HCA attention, mHC, Muon optimizer) and data, but unlike the V3 report it gives no GPU count, GPU-hours or training cost."},{"id":"claude-opus-4-7","model":"Claude Opus 4.7","publisher":"Anthropic","released":"2026-04","activeParamsB":null,"totalParamsB":null,"trainingTokens":"undisclosed","hardware":"AWS and Google Cloud (accelerators not specified)","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":false,"url":"https://www.anthropic.com/news/claude-opus-4-7","notes":"Released April 16, 2026. Anthropic's transparency hub lists cloud computing resources from Amazon Web Services and Google Cloud Platform (with PyTorch, JAX and Triton) but no chip types, counts, tokens or cost."},{"id":"llama-4-maverick","model":"Llama 4 Maverick","publisher":"Meta","released":"2025-04","activeParamsB":17,"totalParamsB":400,"trainingTokens":"~22T","hardware":"NVIDIA H100-80GB","hardwareCount":null,"computeHours":"2.38M GPU-hours","estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":true,"url":"https://github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md","notes":"Released April 5, 2025. Meta's model card discloses 2.38M H100 GPU-hours of pretraining for Maverick (Llama 4 Scout took 5.0M on about 40T tokens). Cluster size and cost are not disclosed."},{"id":"gemini-2.5-pro","model":"Gemini 2.5 Pro","publisher":"Google","released":"2025-03","activeParamsB":null,"totalParamsB":null,"trainingTokens":"undisclosed","hardware":"Google TPUv5p (multiple 8,960-chip pods)","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":false,"url":"https://arxiv.org/abs/2507.06261","notes":"First announced March 25, 2025. The Gemini 2.5 technical report says the family was the first trained on TPUv5p, with synchronous data-parallel training across multiple 8,960-chip pods spread over several datacenters. No parameter, token or cost figures."},{"id":"gpt-4","model":"GPT-4","publisher":"OpenAI","released":"2023-03","activeParamsB":280,"totalParamsB":1760,"trainingTokens":"~13T (estimated)","hardware":"NVIDIA A100","hardwareCount":25000,"computeHours":"~270M GPU-hours (estimated)","estimatedCostMillionUSD":100,"costSource":"estimated","duration":"~3 months","openWeights":false,"url":"https://openai.com/index/gpt-4-research/","notes":"Reverse-engineered numbers from public leaks (1.76T MoE, 280B active). The reference point for \"what did frontier training cost in 2023.\" Made obsolete in compute terms by H100/H200/Blackwell era."},{"id":"llama-3.1-405b","model":"Llama 3.1 405B","publisher":"Meta","released":"2024-07","activeParamsB":405,"totalParamsB":405,"trainingTokens":"15T+","hardware":"NVIDIA H100-80GB","hardwareCount":16000,"computeHours":"30.84M GPU-hours","estimatedCostMillionUSD":60,"costSource":"estimated","duration":"undisclosed (paper analyzes a 54-day pretraining window)","openWeights":true,"url":"https://arxiv.org/abs/2407.21783","notes":"Meta's model card documents 30.84M H100 GPU-hours for the 405B model, and the Llama 3 paper says it trained on up to 16K H100s. The ~$60M cost is an estimate at roughly $2 per GPU-hour, not a Meta disclosure."},{"id":"deepseek-v3","model":"DeepSeek V3","publisher":"DeepSeek","released":"2024-12","activeParamsB":37,"totalParamsB":671,"trainingTokens":"14.8T","hardware":"NVIDIA H800","hardwareCount":2048,"computeHours":"2.788M GPU-hours","estimatedCostMillionUSD":5.576,"costSource":"disclosed","duration":"under two months (pretraining)","openWeights":true,"url":"https://github.com/deepseek-ai/DeepSeek-V3","notes":"The paper reports 2.664M H800 GPU-hours for pretraining plus 119K for context extension and 5K for post-training on a 2,048-GPU cluster. The $5.576M figure assumes $2 per GPU-hour and covers only the final run, excluding prior research and ablations."},{"id":"olmo-2-32b","model":"OLMo 2 32B","publisher":"Allen AI","released":"2025-03","activeParamsB":32,"totalParamsB":32,"trainingTokens":"6.6T","hardware":"NVIDIA H100 (Ai2 Jupiter and Augusta clusters)","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":true,"url":"https://arxiv.org/abs/2501.00656","notes":"Fully open: weights, training data, code, logs and intermediate checkpoints. The report gives 6.6T total tokens for the 32B model (6.06T in pretraining, the rest in mid-training) and describes Ai2's H100 clusters, but not GPU-hours for this run."},{"id":"mistral-large-2","model":"Mistral Large 2","publisher":"Mistral","released":"2024-07","activeParamsB":123,"totalParamsB":123,"trainingTokens":"undisclosed","hardware":"undisclosed","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":true,"url":"https://mistral.ai/news/mistral-large-2407/","notes":"Released July 24, 2024 under the Mistral Research License: weights are downloadable for research and non-commercial use, and commercial self-deployment needs a separate license. Mistral does not disclose tokens, hardware or cost."},{"id":"qwen-2.5-72b","model":"Qwen 2.5 72B","publisher":"Alibaba","released":"2024-09","activeParamsB":72,"totalParamsB":72,"trainingTokens":"18T","hardware":"undisclosed","hardwareCount":null,"computeHours":null,"estimatedCostMillionUSD":null,"costSource":"estimated","duration":"undisclosed","openWeights":true,"url":"https://arxiv.org/abs/2412.15115","notes":"Alibaba's open dense flagship. The Qwen2.5 technical report states 18T pretraining tokens (up from 7T for Qwen2) but gives no hardware, GPU-hours or cost."}]}