Alibaba has detailed a comprehensive AI roadmap covering foundation models, proprietary processors, cloud infrastructure and software designed for deploying AI agents. Unveiled during the annual Apsara Conference, these updates highlight the training of Qwen 4, forward-looking plans for larger Qwen models, the launch of the Zhenwu V900 AI accelerator, and fresh Alibaba Cloud services built to construct and operate AI agents.

To underpin these technical ambitions amid accelerating demand for AI computing, Eddie Wu, CEO of Alibaba Group, outlined a major expansion for the company’s backbone infrastructure. Under this plan, Alibaba Cloud aims to operate more than 20GW of global data centre capacity by 2032.

Qwen roadmap extends to models with trillions of parameters

Alibaba confirmed that Qwen 4 is currently in training, whilst also outlining development pathways for Qwen 4.5 and Qwen 5. The forthcoming models across this flagship series are projected to reach between 5 trillion and 10 trillion parameters, marking a dramatic increase in architectural scale. Alongside these capacity expansions, the company reported measurable progress in recursive self-improvement, a technique where models evaluate feedback from previous attempts to refine subsequent outputs without manual intervention.

Empirical testing highlighted this autonomous refinement during a month of automated runs that covered pipeline design, data validation, experimentation and error diagnosis. Over this period, Qwen3.8-Max completed 33 iterative cycles, lifting its Artificial Analysis score from 40 to 45. In a separate chip-design experiment, the model spent more than 60 hours refining its output and initiated over 10,000 calls to electronic design automation tools. These automated adjustments yielded chip bus modules that reduced chip area by 42% without reducing performance.

Beyond foundational reasoning, the wider Qwen family is expanding its capabilities across translation, speech, audio and image generation. The new Qwen3.8-LiveTranslate model is tailored for simultaneous interpretation, with Alibaba reporting that its latency fell by nearly 20%, dropping from 2.8 seconds to 2.3 seconds. For multimedia production, Qwen-Audio-3.1-TTS-Next can generate dialogue alongside ambient audio directly from a text script to serve audiobooks, film, television, podcasts and games. This multimodal suite also features updated models for automatic speech recognition, text-to-speech and real-time voice interaction, alongside Qwen-Image 3.1, which is scheduled to launch later this year with support for transparent-background generation and image editing.

On mobile devices, the company introduced Qwen Intelligence as a dedicated agent platform aimed at smartphone manufacturers. The framework enables AI phones to execute complex tasks across multiple applications.

Zhenwu V900 expands Alibaba’s AI computing hardware

T-Head, Alibaba’s chip design unit, introduced the Zhenwu V900 processor to handle demanding AI training and inference tasks. The accelerator integrates 216GB of memory with 1,200GB/s of inter-chip bandwidth, alongside broad support for advanced computing formats including FP8 and FP4. According to Alibaba, the V900 yields three times the performance of the earlier Zhenwu M890 released in May. Mass production and a full commercial release are currently scheduled for the first quarter of 2027.

Market uptake for the family is already established across several key industries. According to Alibaba, T-Head’s Zhenwu processors are in active use by more than 650 customers operating in automotive, finance, large language models, embodied intelligence, energy and manufacturing environments.

To assemble these accelerators into massive computational fabrics, the company unveiled an upgraded supernode server that merges the V900 with proprietary networking, storage and controller components. This integrated architecture can support immense clusters containing as many as 500,000 accelerator cards. Alongside accelerator advancements, the processor roadmap incorporates the Yitian 720 and Yitian 730 CPUs, which are both scheduled for release in 2027. The Yitian 720 brings targeted enhancements in single-core performance, core density and energy efficiency over the prior Yitian 710. Advancing the architecture further, the Yitian 730 will stand as the first CPU constructed on T-Head’s proprietary microarchitecture, with Alibaba claiming it delivers up to 40% higher SPECint2017/GHz performance than its Yitian 710 predecessor.

Cloud infrastructure adds tools for enterprise AI agents

Alibaba Cloud is organising its broader AI infrastructure around three foundational pillars, specifically model computing, agent deployment and the data required to give agents operational context. Within this framework, AI Native Cloud provides the underlying hardware and software needed to train and run complex models. Its integrated Platform for AI brings together inference, caching, sample replay and training services, allowing teams to complete the post-training of Qwen models within five days. Supporting this computational throughput, the Cloud Parallel File Storage system manages the massive volumes of data required during AI model training, which Alibaba claims can lower enterprise AI storage expenditure by 69%. Networking capacity has similarly scaled via the HPN 8.0 Pro architecture, supplying 100 petabits of aggregate bandwidth and accommodating over 130,000 network ports operating at 800G speeds within a single cluster.

At the software layer, the company introduced AgentCore to assist organisations in building, running and orchestrating AI agents across their complete operational lifecycle. The environment features built-in controls for human-agent collaboration along with telemetry tools to monitor live performance. To safeguard these autonomous workflows, the companion Agent Security Center adds dedicated compliance management paired with real-time threat detection.

To ensure that software agents retain pertinent knowledge, the Context Engine equips them with up-to-date information and sustained memory. Its Agent Context service bridges disparate corporate assets by connecting documents, business systems, chat records and multimodal data. This arrangement allows agents to preserve vital context from earlier interactions and exchange knowledge across continuous workflows. In practical deployment, Alibaba reports that Agent Context can cut token consumption by up to 67% across operational tasks such as customer service, software engineering and data analytics.

Data handling has been further consolidated through upgrades to OpenLake, which now processes structured, semi-structured, unstructured, vector and streaming data within a unified data lakehouse environment. Alibaba claims that these architectural enhancements can reduce total operating costs by 38% while cutting query response times by 40%.

Share