[ INTEL_NODE_32848 ] · PRIORITY: 8.8/10

Legacy Beast Awakens: IBM AC922 Hits 7,300+ tk/s Prefill on Qwen via Optimized Strata Fork

●  PUBLISHED: · SOURCE: Reddit LocalLLaMA →
[ DATA_STREAM_START ]

Core Event

A developer has successfully revitalized the 2018-era IBM AC922 server by forking Strata (Opus 5.5) to optimize LLM inference. Running Qwen3.8-FN (UD-Q4_K_XL) on a dual POWER9 CPU and 4x NVIDIA Tesla V100 setup, the system achieved a blistering prefill speed of 7,357 tk/s and a decode speed of 113 tk/s. The breakthrough leverages the machine’s unique CPU-to-GPU NVLink 2.0 interconnect, which offers 150GB/s of bandwidth and robust Unified Memory support.

  • ▶ Bypassing x86 Constraints: While standard llama.cpp struggles on POWER9 due to the absence of AVX instructions, this custom implementation optimizes for the architecture’s specific vector units and memory layout.
  • ▶ Interconnect Supremacy: The 150GB/s CPU-GPU bandwidth allows the system to treat system RAM and VRAM as a more cohesive pool, drastically accelerating the prefill phase compared to modern PCIe-based consumer setups.

Bagua Insight

This project highlights a critical industry oversight: the “Compute Bottleneck” is often actually an “Interconnect Bottleneck.” While the world chases H100 clusters, this experiment proves that high-bandwidth legacy hardware can still punch significantly above its weight class in the GenAI era. The AC922’s ability to sustain 7,300+ tk/s prefill makes it a formidable candidate for RAG (Retrieval-Augmented Generation) workloads, where ingesting massive contexts quickly is more vital than raw token generation speed. It serves as a masterclass in hardware-software co-design, proving that software tailored to hardware topology can outperform generic solutions on much newer silicon.

Actionable Advice

Enterprises and research labs sitting on legacy HPC assets (specifically POWER9/V100 nodes) should reconsider decommissioning. Instead: 1. Audit these clusters for high-throughput RAG pipelines where prefill latency is the primary bottleneck; 2. Invest in custom inference stacks that bypass x86-centric limitations to unlock latent TFLOPS in non-standard enterprise architectures.

[ DATA_STREAM_END ]
[ ORIGINAL_SOURCE ]
READ_ORIGINAL →
[ 02 ] RELATED_INTEL