News Article · Aug 2, 2026 at 11:42 PM
3 min read 0
Member
AMD MI355X Beats NVIDIA B300 on AI Inference Cost per Token for Kimi K3 Model
Cloud #AI inference #Kimi K3 #AMD MI355X #NVIDIA B300 #HBM memory #Majestic Labs #LPDDR6 #Arm servers #performance per dollar

AMD MI355X Beats NVIDIA B300 on AI Inference Cost per Token for Kimi K3 Model

Wafer's benchmark of the 2.8T-parameter Kimi K3 model shows AMD's MI355X GPU delivers 2.4x lower cost per GPU than NVIDIA's B300 while achieving competitive throughput. Startups like Majestic Labs are exploring Arm-based servers with LPDDR6 memory to bypass HBM bottlenecks.

The 2.8-trillion-parameter Kimi K3 model, which requires more than 1.5 TB of VRAM before KV cache allocation, cannot fit on a single node of NVIDIA B200 GPUs. Wafer, an AI inference startup, has benchmarked the model on AMD's MI355X GPU and found it achieves higher performance per dollar than NVIDIA's B300, at 48 tok/s per dollar per GPU-hour versus 33 tok/s for the B300.

On a 1,024-token input, 400-token output benchmark, the MI355X node (8 GPUs) delivered 952 tok/s aggregate throughput and 118 tok/s single-stream decode. That compares to 498 tok/s aggregate across two B200 nodes (16 GPUs, 249 tok/s per node) and 1,568 tok/s aggregate on a B300 node (8 GPUs with DCP). The B300 still wins on raw throughput, but at 2.4 times the price per GPU, the MI355X delivers roughly 1.45 times the performance per dollar.

Memory capacity as a competitive moat

Kimi K3's size creates a memory capacity problem that favors AMD's MI355X, which has 288 GB of HBM3 per GPU, matching the B300's capacity and exceeding the B200's 192 GB. Wafer notes that the MI355X is about 2.4 times cheaper per GPU than the B300 and 1.7 times cheaper than the B200, making it the only non-NVIDIA GPU that can fit the model on a single node. The company overcame software compatibility issues, including a missing top-k renorm kernel in ROCm, by implementing a simple PyTorch function that fixed the speculative decoding pipeline, boosting single-stream performance by 2.2 times.

  • MI355X: ~$2.50 per GPU-hour, 952 tok/s aggregate, 48 tok/s per dollar
  • B300: ~$6.00 per GPU-hour, 1,568 tok/s aggregate, 33 tok/s per dollar
  • B200: ~$4.25 per GPU-hour, 498 tok/s across two nodes, 7 tok/s per dollar
  • Kimi K3 requires 8 GPUs with 288 GB each to fit weights plus 1M-token KV cache

Beyond GPU: alternative memory architectures

Startup Majestic Labs has unveiled an Arm-based AI server that replaces GPU HBM with unified LPDDR6 memory, claiming capacities up to 128 TB at a fraction of the cost of HBM. The approach uses many Arm cores to process data locally, avoiding the memory wall that limits GPU-based systems. Separately, researchers have published work on persistent state machines that use INT4 in-memory cells for LLM attention, further reducing memory bandwidth requirements.

These developments point to a diversification of AI hardware. While NVIDIA's B300 remains the performance leader, the combination of model size growth and HBM cost pressures is creating room for alternatives. Wafer plans to release benchmarks for the MI355X on other large models, including DeepSeek V4-Pro and GLM5.2, as the industry weighs the tradeoffs between raw throughput, memory capacity, and total cost of ownership.

Fact check

  • AMD MI355X costs ~2.4x less per GPU than NVIDIA B300

    reported · source

  • MI355X delivers 48 tok/s per dollar per GPU-hour vs 33 tok/s for B300

    reported · source

  • Kimi K3 has 2.8 trillion parameters and requires over 1.5 TB of VRAM

    reported · source

  • Majestic Labs unveiled an Arm-based AI server with up to 128 TB of LPDDR6 memory

    reported · source

Source reporting (3)

0 Comments

No comments yet

Be the first to share your thoughts on this article.

Join the conversation

You need to be registered and logged in to comment on blog articles.

Who Is Online

In total there are 52 users online: 0 registered, 44 guests and 8 bots.

Most users ever online was 9,867 on 30 Jul 2026, 2:30 am.

Bots: AhrefsBot Baiduspider Bingbot Googlebot Other Bot Other Crawler PetalBot SemrushBot

Users active in the past 15 minutes. Total registered members: 374