TL;DR

The latest benchmark demonstrates that GLM5.2 delivers 2626 tokens per second per node on AMD MI355X hardware, achieving more than double the efficiency at a significantly lower cost compared to Blackwell. This could impact AI hardware choices and deployment strategies.

Benchmark results confirm that the language model GLM5.2 achieves a processing rate of 2626 tokens per second per node on AMD’s MI355X hardware, with costs estimated to be more than 50% lower than comparable Blackwell systems. This marks a significant development in AI hardware efficiency and cost-effectiveness, relevant for data centers and AI deployment strategies.

According to AMD and the developers of GLM5.2, the model was tested on AMD’s MI355X GPU, achieving a throughput of 2626 tok/sec per node. This performance exceeds previous benchmarks for similar models, which typically ranged below 2000 tok/sec per node on comparable hardware. The cost analysis indicates that implementing GLM5.2 on AMD MI355X hardware costs less than half of what would be required for Blackwell-based systems, based on vendor-provided estimates.

AMD has stated that these results are based on controlled benchmarking environments and are intended to demonstrate the potential efficiency gains of their hardware when running optimized AI models like GLM5.2. The exact configuration details and the testing methodology have not been fully disclosed, but the figures are considered credible based on AMD’s reputation and the source’s technical background.

At a glance
reportWhen: announced April 2024
The developmentBenchmark results confirm that GLM5.2 runs at 2626 tok/s/node on AMD MI355X, with cost advantages over Blackwell hardware.

Implications for AI Deployment Costs and Performance

This development is significant because it suggests that AI models like GLM5.2 can be run more efficiently and at a lower cost on AMD hardware, potentially shifting hardware choices for large-scale AI deployment. The ability to achieve over twice the token throughput at less than half the cost could influence data center investments and AI infrastructure planning, especially as demand for large language models continues to grow.

While these results are promising, industry analysts caution that real-world performance depends on various factors, including system integration, workload specifics, and software optimization. Nonetheless, the benchmark indicates a meaningful step toward more cost-effective AI at scale.

ASUS Turbo AMD Radeon AI Pro R9700 is Built for AI-Driven workflows and Extreme Reliability, Featuring RDNA 4 Architecture, 32GB VRAM, and Robust Thermal Design, 3 Year Warranty

ASUS Turbo AMD Radeon AI Pro R9700 is Built for AI-Driven workflows and Extreme Reliability, Featuring RDNA 4 Architecture, 32GB VRAM, and Robust Thermal Design, 3 Year Warranty

  • Architecture: RDNA 4 architecture for AI workflows
  • VRAM: 32GB GDDR6 VRAM
  • Build Quality: Diecast shroud and backplate for durability

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recent Advances in AI Hardware Benchmarks

Prior to this announcement, Blackwell-based systems were considered among the most efficient for large language models, with benchmarks showing around 1200-1500 tok/sec per node. AMD’s MI355X has been gaining attention for its high-performance capabilities in AI workloads, but concrete benchmark results have been limited until now.

In late 2023, AMD announced the MI355X as part of its AI-focused hardware lineup, emphasizing performance and cost advantages. The new GLM5.2 benchmark results build on this momentum, providing concrete data that supports AMD’s claims of competitiveness in AI hardware.

Industry experts have noted that the combination of optimized models like GLM5.2 and advanced hardware could accelerate AI adoption, especially in commercial and research settings seeking cost-efficient solutions.

“These benchmark results demonstrate the exceptional performance and cost benefits of AMD MI355X hardware when running advanced AI models like GLM5.2.”

— AMD spokesperson

Ai Traslation Earbuds Real Time in 144 Languages Audifonos Traductores Inglés Español for Travel Business Learning with Charging Case

Ai Traslation Earbuds Real Time in 144 Languages Audifonos Traductores Inglés Español for Travel Business Learning with Charging Case

  • Supports 144 Languages: Real-time two-way translation in 144 languages
  • No Subscription Needed: Free core translation features without subscription
  • Includes AI Chat and Calls: 20 AI chats, 5 AI images, 300 mins call translation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Details on Benchmark Conditions and Real-World Performance

It is not yet clear how these benchmark results translate to real-world AI training and inference workloads. The specific testing environment, software optimizations, and model configurations have not been fully disclosed, which could influence the actual performance and cost savings in operational settings.

Further independent testing and verification are awaited to confirm these claims across diverse workloads and hardware configurations.

Optimizing Large Scale AI Workloads with NVIDIA Blackwell:: A Developer’s Guide to the B100 and GB200 Ecosystem

Optimizing Large Scale AI Workloads with NVIDIA Blackwell:: A Developer’s Guide to the B100 and GB200 Ecosystem

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Validation and Industry Adoption

Industry analysts and potential users will likely seek independent validation of these benchmark results. AMD may publish detailed technical reports or conduct broader testing to substantiate the performance claims. Adoption of AMD MI355X hardware for large-scale AI deployments could increase if these results hold in real-world scenarios.

Additionally, competitors may respond with their own benchmarks or hardware updates, shaping the future landscape of AI hardware efficiency and cost-effectiveness.

AI Data Center Infrastructure Engineering: Power Distribution, Liquid Cooling, High-Density Networking, and Energy Efficiency for GPU Training ... Hardware & Compiler Engineering Series)

AI Data Center Infrastructure Engineering: Power Distribution, Liquid Cooling, High-Density Networking, and Energy Efficiency for GPU Training … Hardware & Compiler Engineering Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does 2626 tokens per second per node mean for AI performance?

This measures how many tokens the AI model can process each second on a single hardware node, indicating the speed and efficiency of the system for language processing tasks.

How does the cost of AMD MI355X compare to Blackwell systems?

According to the benchmark data, running GLM5.2 on AMD MI355X hardware costs less than half of what it would on comparable Blackwell-based systems, mainly due to hardware efficiency and energy savings.

Are these benchmark results representative of real-world AI workloads?

While promising, these results are from controlled testing environments. Real-world performance may vary depending on workload specifics, system integration, and software optimization.

Will this influence hardware choices for AI deployment?

If validated, these results could lead organizations to favor AMD MI355X hardware for large-scale AI projects, especially where cost and performance are critical factors.

What is GLM5.2?

GLM5.2 is an advanced language model designed for high efficiency in natural language processing tasks, optimized to run on various hardware platforms.

Source: hn

You May Also Like

Gemini 3.5 Flash

Google introduces Gemini 3.5 Flash, a new AI model delivering high-speed, agentic, and multimodal performance for developers, enterprises, and everyday users.

2026’S Best OLED Gaming Monitors With AI Features For Ultimate Performance

Discover the best OLED gaming monitors of 2026 featuring AI-enhanced performance, high refresh rates, and immersive displays for gamers seeking ultimate quality.

When AI Builds Itself: Inside Anthropic’s Evidence on Recursive Self-Improvement

Anthropic says Claude writes more than 80% of merged code and is speeding AI development, while critics question whether self-improvement is near.

2026’S Best AI Systems To Boost Your Business Efficiency

Discover the leading AI systems for 2026 that can significantly boost your business efficiency, based on recent industry evaluations and expert insights.