[ INTEL_NODE_31768 ] · PRIORITY: 9.2/10

GPU-Free Future? Alibaba’s XuanTie C950 RISC-V CPU Hits 30 TPS on 27B Qwen Model

  PUBLISHED: · SOURCE: Reddit LocalLLaMA →
[ DATA_STREAM_START ]

Alibaba’s chip division, T-Head, has demonstrated a significant breakthrough in AI inference performance using its XuanTie C950 RISC-V processor. The CPU achieved a sustained inference speed of 30 tokens per second (tps) while running the Qwen-3.8 27B parameter model. This benchmark signals that RISC-V is no longer just for low-power IoT, but a serious contender in the high-performance generative AI landscape.

  • Performance Milestone: Achieving 30 tps on a 27B model is a high-water mark for CPU-based inference, effectively rivaling dedicated mid-range AI accelerators for localized workloads.
  • Architectural Prowess: The C950 leverages advanced RISC-V Vector (RVV) extensions and optimized matrix math units to bypass the traditional bottlenecks associated with general-purpose CPUs in Transformer-based tasks.
  • Strategic Decoupling: By vertically integrating its own silicon (XuanTie) with its proprietary LLM (Qwen), Alibaba is showcasing a viable path for high-performance AI that is independent of the x86/ARM duopoly and high-end GPU dependencies.

Bagua Insight

This is a watershed moment for the RISC-V ecosystem. The 27B parameter class is widely considered the “sweet spot” for enterprise-grade local LLMs—powerful enough for complex reasoning but demanding in terms of memory bandwidth and compute. Alibaba’s ability to hit 30 tps on a CPU suggests that the “GPU tax” for edge AI and private cloud deployments could soon be optional. This isn’t just about raw speed; it’s about democratizing high-quality AI by making it run efficiently on versatile, cost-effective RISC-V hardware. Alibaba is effectively building a full-stack hedge against global GPU supply chain volatility.

Actionable Advice

Infrastructure leads should re-evaluate RISC-V as a cost-effective alternative for inference-heavy workloads, particularly in edge computing environments where power efficiency and TCO are critical. AI software teams should prioritize mastering RVV-compatible kernels and optimization libraries to future-proof their deployment stacks against a more fragmented and competitive hardware landscape.

[ DATA_STREAM_END ]
[ ORIGINAL_SOURCE ]
READ_ORIGINAL →
[ 02 ] RELATED_INTEL