Product Information

AI-RAN: Shared GPU Infrastructure for 5G RAN and Edge AI NEWS DETAIL

Current Position:Home > News and Insights
Category: News and Insights Author: Zhongke Xinyuan Content Reviewer: Zhongke Xinyuan Review Published: 2025-02-06 Updated: 2026-07-22 Source: Existing page; verify sources
AI-RAN: Shared GPU Infrastructure for 5G RAN and Edge AI

AI-RAN is an approach to running 5G radio access network (RAN) functions and AI workloads on shared GPU-based infrastructure. The source record describes SoftBank’s outdoor trial in Fujisawa, Kanagawa, Japan, using NVIDIA accelerated hardware and NVIDIA Aerial software. For operators, the practical question is whether existing centralized and distributed network sites can support AI inference while preserving required RAN performance. That answer depends on workload isolation, orchestration behavior, network design, and validation in the operator’s own spectrum, radio, and traffic conditions.

The problem AI-RAN is intended to address

Conventional CPU- or ASIC-based RAN systems are designed primarily for radio workloads. The source positions AI-RAN as a way to use a common GPU-based compute foundation for both wireless processing and AI workloads, rather than deploying separate infrastructure for each purpose.

This model is relevant where AI inference must interact with local devices or data streams over mobile connectivity. Examples demonstrated in the described trial include 5G-connected remote support for an autonomous vehicle, multimodal factory applications, and robotics. These scenarios may benefit from placing inference closer to connected endpoints, but latency, privacy, resilience, and data-governance requirements must be evaluated for each application.

Architecture described in the SoftBank trial

The trial implementation combined a software-defined 5G vRAN stack with AI workloads. According to the source, it used NVIDIA GH200 systems, NVIDIA BlueField-3 NIC/DPU hardware, and Spectrum-X networking for fronthaul and backhaul. It integrated 20 radio units, a 5G core network, and 100 mobile UEs.

  • NVIDIA Aerial CUDA-accelerated RAN libraries were used for Layer 1 functions, including channel mapping, channel estimation, modulation, and forward error correction.
  • Fujitsu software supported Layer 2 functions.
  • Red Hat OpenShift Container Platform provided the container virtualization layer.
  • A SoftBank end-to-end AI and RAN orchestrator allocated workloads according to demand and available capacity.

The source also describes NVIDIA Multi-Instance GPU (MIG) as a mechanism for partitioning a GPU into independent instances with dedicated resources. In principle, this enables allocations for AI and RAN workloads to be separated spatially, while an orchestration policy can also allocate resources according to time or workload priority.

Capabilities and deployment scenarios

The reported field trial demonstrated concurrent AI and RAN processing with static resource allocation. The source states that a single GH200 server, in RAN-only mode, processed 20 5G cells using 100 MHz bandwidth, with a reported 1.3 Gbps peak downlink result per cell under ideal conditions and 816 Mbps described as commercial-grade availability in the outdoor deployment.

These figures are trial-specific and should not be treated as deployment guarantees. Cell capacity depends on radio configuration, spectrum, antenna configuration, user distribution, propagation conditions, software release, transport design, and service-level objectives.

Suitable evaluation scenarios include edge inference for video, audio, and sensor streams; enterprise applications requiring local data handling; and robotics or connected-vehicle workflows where application response time is material. A cloud-only design may remain appropriate when local inference capacity, operations complexity, or site economics do not justify distributed AI infrastructure.

How operators should evaluate AI-RAN

  1. Define RAN protection requirements. Establish radio workload priority, required throughput, latency, availability, and behavior during traffic peaks before assigning any compute capacity to AI.
  2. Select a narrow AI workload. Begin with an inference use case that has measurable latency, data-locality, and operational requirements.
  3. Test isolation and orchestration. Validate GPU partitioning, workload admission, failover, observability, and the effect of AI load changes on RAN key performance indicators.
  4. Model site-specific economics. Include hardware, software, power, cooling, transport, operations, AI utilization, and local regulatory constraints in the total-cost model.
  5. Verify the complete bill of materials. Confirm supported hardware, software versions, interfaces, radio configurations, and licensing in dated official documentation.

Evidence boundaries and commercial claims

The source reports further economic and energy comparisons for a GB200-NVL2-based NVIDIA Aerial RAN Computer-1 target platform, including assumptions for a 600-cell Tokyo-area model and Japanese electricity costs. It also reports performance claims for specific AI workloads. Those results rely on the stated platform, workload, utilization mix, and local cost assumptions; they are not sufficient evidence for another operator’s business case.

The source says SoftBank targeted commercial release of its own AI-RAN product in 2026 and planned a reference kit. Readers should verify the status, scope, availability, supported configurations, and contractual terms through dated official SoftBank and NVIDIA documentation.

FAQ

Can AI and RAN workloads run on the same GPU infrastructure?

The source describes concurrent AI and RAN processing on GH200 systems, using static allocation in the reported trial and MIG-based partitioning as part of the resource model. Production suitability requires testing under peak radio traffic, AI demand spikes, failure conditions, and the operator’s own RAN performance objectives.

Does the trial prove that AI-RAN will reduce cost or create revenue?

No. The source presents modeled economics based on selected assumptions, including Japanese local electricity costs, a five-year TCO period, and estimated AI-compute revenue. These inputs must be independently validated against local energy prices, utilization, monetization demand, deployment scale, and operating costs.

Conclusion

AI-RAN presents a shared-infrastructure path for combining software-defined 5G RAN and edge AI inference. The SoftBank trial provides an implementation example and reported field evidence, but operators should treat it as a starting point for technical and commercial validation rather than a universal deployment result. The decision should be based on protected RAN performance, workload isolation, real AI demand, and a site-specific total-cost model.

After reviewing AI-RAN: Shared GPU Infrastructure for 5G RAN and Edge AI, continue with NVIDIA products and networking solutions for related evaluation paths.