Core EventA developer has successfully revitalized the 2018-era IBM AC922 server by forking Strata (Opus 5.5) to optimize LLM inference. Running Qwen3.8-FN (UD-Q4_K_XL) on a dual POWER9 CPU and 4x NVIDIA Tesla V100 setup, the system achieved a blistering prefill speed of 7,357 tk/s and a decode speed of 113 tk/s. The breakthrough leverages the machine's unique CPU-to-GPU NVLink 2.0 interconnect, which offers 150GB/s of bandwidth and robust Unified Memory support.▶ Bypassing x86 Constraints: While standard llama.cpp struggles on POWER9 due to the absence of AVX instructions, this custom implementation optimizes for the architecture's specific vector units and memory layout.▶ Interconnect Supremacy: The 150GB/s CPU-GPU bandwidth allows the system to treat system RAM and VRAM as a more cohesive pool, drastically accelerating the prefill phase compared to modern PCIe-based consumer setups.Bagua InsightThis project highlights a critical industry oversight: the "Compute Bottleneck" is often actually an "Interconnect Bottleneck." While the world chases H100 clusters, this experiment proves that high-bandwidth legacy hardware can still punch significantly above its weight class in the GenAI era. The AC922’s ability to sustain 7,300+ tk/s prefill makes it a formidable candidate for RAG (Retrieval-Augmented Generation) workloads, where ingesting massive contexts quickly is more vital than raw token generation speed. It serves as a masterclass in hardware-software co-design, proving that software tailored to hardware topology can outperform generic solutions on much newer silicon.Actionable AdviceEnterprises and research labs sitting on legacy HPC assets (specifically POWER9/V100 nodes) should reconsider decommissioning. Instead: 1. Audit these clusters for high-throughput RAG pipelines where prefill latency is the primary bottleneck; 2. Invest in custom inference stacks that bypass x86-centric limitations to unlock latent TFLOPS in non-standard enterprise architectures.
SOURCE: REDDIT LOCALLLAMA // UPLINK_STABLE