AMD MI355X Beats NVIDIA B300 on AI Inference Cost per Token for Kimi K3 Model
Wafer's benchmark of the 2.8T-parameter Kimi K3 model shows AMD's MI355X GPU delivers 2.4x lower cost per GPU than NVIDIA's B300 while achieving competitive throughput. Startups like Majestic Labs are exploring Arm-based servers with LPDDR6 memory to bypass HBM bottlenecks.
The 2.8-trillion-parameter Kimi K3 model, which requires more than 1.5 TB of VRAM before KV cache allocation, cannot fit on a single node of NVIDIA B200 GPUs. Wafer, an AI inference startup, has benchmarked the model on AMD's MI355X GPU and found it achieves higher performance per dollar than NVIDIA's B300, at 48 tok/s per dollar per GPU-hour versus 33 tok/s for the B300.
On a 1,024-token input, 400-token output benchmark, the MI355X node (8 GPUs) delivered 952 tok/s aggregate throughput and 118 tok/s single-stream decode. That compares to 498 tok/s aggregate across two B200 nodes (16 GPUs, 249 tok/s per node) and 1,568 tok/s aggregate on a B300 node (8 GPUs with DCP). The B300 still wins on raw throughput, but at 2.4 times the price per GPU, the MI355X delivers roughly 1.45 times the performance per dollar.
Memory capacity as a competitive moat
Kimi K3's size creates a memory capacity problem that favors AMD's MI355X, which has 288 GB of HBM3 per GPU, matching the B300's capacity and exceeding the B200's 192 GB. Wafer notes that the MI355X is about 2.4 times cheaper per GPU than the B300 and 1.7 times cheaper than the B200, making it the only non-NVIDIA GPU that can fit the model on a single node. The company overcame software compatibility issues, including a missing top-k renorm kernel in ROCm, by implementing a simple PyTorch function that fixed the speculative decoding pipeline, boosting single-stream performance by 2.2 times.
- MI355X: ~$2.50 per GPU-hour, 952 tok/s aggregate, 48 tok/s per dollar
- B300: ~$6.00 per GPU-hour, 1,568 tok/s aggregate, 33 tok/s per dollar
- B200: ~$4.25 per GPU-hour, 498 tok/s across two nodes, 7 tok/s per dollar
- Kimi K3 requires 8 GPUs with 288 GB each to fit weights plus 1M-token KV cache
Beyond GPU: alternative memory architectures
Startup Majestic Labs has unveiled an Arm-based AI server that replaces GPU HBM with unified LPDDR6 memory, claiming capacities up to 128 TB at a fraction of the cost of HBM. The approach uses many Arm cores to process data locally, avoiding the memory wall that limits GPU-based systems. Separately, researchers have published work on persistent state machines that use INT4 in-memory cells for LLM attention, further reducing memory bandwidth requirements.
These developments point to a diversification of AI hardware. While NVIDIA's B300 remains the performance leader, the combination of model size growth and HBM cost pressures is creating room for alternatives. Wafer plans to release benchmarks for the MI355X on other large models, including DeepSeek V4-Pro and GLM5.2, as the industry weighs the tradeoffs between raw throughput, memory capacity, and total cost of ownership.
Fact check
-
AMD MI355X costs ~2.4x less per GPU than NVIDIA B300
reported · source
-
MI355X delivers 48 tok/s per dollar per GPU-hour vs 33 tok/s for B300
reported · source
-
Kimi K3 has 2.8 trillion parameters and requires over 1.5 TB of VRAM
reported · source
-
Majestic Labs unveiled an Arm-based AI server with up to 128 TB of LPDDR6 memory
reported · source
Source reporting (3)
- Hacker News Front Page · Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
- TechRadar Pro · Startup swaps costly AI GPUs for Arm cores and up to 128TB of 'cheap' LPDDR6 RAM instead of expensive HBM to smash through the memory wall
- Hacker News Front Page · Persistent State Machines: LLM Attention with INT4 In-Memory Cells
Join the conversation
You need to be registered and logged in to comment on blog articles.
Related Articles
Microsoft Azure Surpasses $100B Annual Revenue as AI Investments Deliver Mixed Returns
Jul 30, 2026
Microsoft's open-weight AI push is an Azure play as 25 companies urge US against restrictions
Jul 25, 2026
Anthropic launches Claude Opus 5, claims near-Fable 5 performance at half the token price
Jul 24, 2026
0 Comments
No comments yet
Be the first to share your thoughts on this article.