
NVIDIA and SoftBank Group outlined a Japan-focused AI infrastructure program spanning AI supercomputing, AI-RAN, and a localized AI service marketplace. Based on the December 6, 2024 source record, the program combines NVIDIA Blackwell and Grace Blackwell platform plans, NVIDIA AI Enterprise software, NVIDIA Quantum-2 InfiniBand networking, and NVIDIA AI Aerial technology. For telecom and enterprise teams, the practical question is whether shared infrastructure can support both 5G network functions and AI inference without compromising required network performance; the source describes trials and plans, not a complete production specification or deployment guarantee.
The problem the program addresses
SoftBank Group’s stated direction is to expand domestic AI computing capacity while finding ways for telecom infrastructure to support AI workloads. Conventional mobile-network infrastructure is principally operated as a network cost base. The AI-RAN approach described in the source is intended to let one infrastructure environment run radio access network workloads and AI workloads, using surplus capacity for inference when available.
The same program also addresses demand for localized AI computing in Japan. SoftBank said it plans to create an AI Marketplace with NVIDIA AI Enterprise, supporting AI training and edge AI inference. The source does not define the marketplace service levels, supported model frameworks, tenant isolation model, data-residency terms, or commercial availability. Those items require verification in dated SoftBank and NVIDIA documentation.
Capabilities described in the source
- SoftBank plans to build a NVIDIA SuperPOD supercomputer using NVIDIA systems, NVIDIA AI Enterprise software, and NVIDIA Quantum-2 InfiniBand networking for large language model development.
- SoftBank stated that it is using the NVIDIA Blackwell platform for an AI supercomputer and plans to use the NVIDIA Grace Blackwell platform for a subsequent supercomputer.
- The AI-RAN trial combines AI and 5G workloads through NVIDIA AI Aerial accelerated computing technology and a software-defined 5G radio stack optimized for NVIDIA AI computing platforms.
- The source identifies NVIDIA Aerial CUDA-accelerated RAN libraries as part of the enhanced Layer 1 software and states SoftBank plans to integrate NVIDIA Aerial RAN Computer-1 into future solutions.
- SoftBank demonstrated AI inference use cases in its trial, including remote assistance for autonomous vehicles, robot control, and edge multimodal retrieval-augmented generation.
- SoftBank describes an internally developed orchestrator intended to coordinate AI and RAN workloads and schedule external inference work to AI-RAN servers when computing resources are available.
Where this architecture may fit
This approach is most relevant to mobile operators evaluating distributed inference close to radio-network infrastructure, and to organizations that need lower-latency inference near operational sites. The source frames potential use cases around transport, robotics, healthcare, and telecom services. It may also be relevant where an operator wants an orchestration layer to place inference jobs only when network capacity permits.
Suitability depends on workload behavior. Latency-sensitive inference, variable demand, and geographically distributed service points can make edge placement attractive. However, training workloads, predictable 5G peak periods, GPU partitioning, backhaul design, and operational ownership must be evaluated separately. The record does not provide server configurations, GPU counts, radio bands, geographic trial scope, latency measurements, utilization thresholds, or availability targets.
Evaluation path for telecom teams
- Define protected RAN performance requirements first, including peak-load behavior, 5G service metrics, and escalation rules when AI demand conflicts with radio workloads.
- Request a complete architecture and BOM covering the NVIDIA platform components, RAN software, orchestration interfaces, networking, storage, and management dependencies.
- Test representative AI inference workloads alongside realistic radio traffic, including overload, failover, maintenance, and capacity-reclamation scenarios.
- Validate how the orchestrator admits, pauses, migrates, or rejects inference jobs when RAN capacity changes.
- Review data handling, access control, model operations, service observability, and local regulatory requirements before exposing inference capacity to external users.
Evidence limits and operational tradeoffs
The source reports that SoftBank conducted outdoor trials and achieved telecom-grade 5G performance while running AI inference with excess capacity. It also contains revenue, return, and power-reduction estimates. These figures should not be treated as universal procurement outcomes: they depend on deployment assumptions that are not fully specified in the supplied record. Buyers should validate performance, power, capacity, economics, and compatibility through dated official documentation, a complete SKU/BOM, and a project-specific proof of concept.
AI-RAN also introduces an operational tradeoff: greater infrastructure utilization may increase scheduling and assurance complexity. A design must demonstrate that AI workloads are governed by RAN priorities, rather than assuming that co-location alone delivers usable excess capacity.
FAQ
Is the AI-RAN solution described as a general-purpose replacement for every 5G deployment?
No. The source describes SoftBank’s trial work and future solution plans. It does not establish universal compatibility, a deployment timeline, or equivalent results across radio environments. Operators should assess their own RAN architecture, traffic profile, hardware configuration, and support model.
What should enterprises verify before using a localized AI inference service?
Verify the actual service architecture, data location and handling terms, supported APIs and models, latency and capacity commitments, security controls, pricing, and incident responsibilities. The source states an intention to build an AI Marketplace, but does not provide those service details.
Conclusion
The NVIDIA and SoftBank program presents AI-RAN as a path for combining AI inference and 5G infrastructure while expanding Japanese AI computing capacity. The source supports the existence of trials, planned platform adoption, and intended orchestration and marketplace capabilities. Production decisions should rest on documented configurations and measured results in the target network, not on trial descriptions or forecast economics alone.
After reviewing NVIDIA and SoftBank AI-RAN Plans for Japan, continue with NVIDIA products and networking solutions for related evaluation paths.

WeChat
Profile