Pinterest has built a shared artificial intelligence infrastructure alongside NVIDIA to power applications that operate across both images and language. This unified architecture supports critical user-facing services across the network, including visual search, content understanding and digital safety systems.
Underpinning the deployment is a combination of NVIDIA Blackwell GPUs and NVIDIA Dynamo integrated with Pinterest’s proprietary visual embeddings. By establishing this centralised technical baseline, engineering groups gain a common foundation to construct and deploy multimodal features efficiently. This coordinated structure removes operational fragmentation, allowing teams to deliver advanced capabilities without the burden of maintaining isolated infrastructure for every distinct tool.
Pinterest expands the visual context available to AI
Pinterest processes more than 80 billion searches every month, generating essential signals for discovery and commerce. The updated infrastructure is tailored to process substantially more visual information within each individual request.
This expanded capability is particularly visible in Pinterest Assistant, which now handles 25 times more visual context per query. Rather than operating on limited visual snapshots, the service draws upon far broader image data when interpreting complex requests. This richer intake ensures that queries relying simultaneously on pictures and written language receive thorough, context-aware analysis.
In benchmark testing, Pinterest documented dramatic performance gains across its serving stack. The company recorded average improvements of about 85 times in response startup alongside a 7.3 times improvement in overall latency.
Much of this efficiency stems from shifting towards precomputed visual representations rather than repeatedly parsing raw imagery. The strategy curbs redundant image processing at scale, providing crucial support as Pinterest broadens its adoption of vision-language models across the business. These multimodal frameworks are vital for products such as Pinterest Assistant, where visual content and written prompts must be evaluated together during a single computational cycle.
“Building the next generation of AI-powered discovery means investing in infrastructure that can keep up with the scale and complexity of Pinterest,” said Kartik Paramasivam, Chief Architect at Pinterest. “Our collaboration with NVIDIA helps us deliver faster, smarter and more personalised experiences for the hundreds of millions of people who use Pinterest.”
One platform supports several AI workloads
The shared architecture was originally created for Pinterest Intelligence, the central AI layer powering the platform. Today, that remit has expanded to support ranking, platform safety, optical character recognition, signal generation and future artificial intelligence experiences.
To achieve this breadth, Pinterest combines NVIDIA Blackwell GPUs and NVIDIA Dynamo with open-source models, proprietary in-house technology and NVIDIA’s broader AI software stack. Unifying these diverse resources into a single underlying framework fundamentally alters engineering workflows. Product teams across the business can build and launch multimodal features collaboratively, removing the complexity and overhead of managing distinct infrastructure stacks for every service.
“Pinterest is transforming visual discovery with AI, helping hundreds of millions of people find inspiration through more intelligent and personalised experiences,” said Ujval Kapasi, vice president, AI & HPC Frameworks and Libraries at NVIDIA. “By building its multimodal AI infrastructure on NVIDIA accelerated computing and software, Pinterest can bring the next generation of visual and conversational AI experiences to its users at scale.”
This milestone represents the culmination of nearly five years of engineering collaboration between Pinterest and NVIDIA. Today, the resulting technical footprint actively employs more than 14,000 NVIDIA GPUs across the company’s computing network.
Standardising this foundational layer provides development teams with a coherent, shared environment for launching both visual and conversational capabilities. Rather than building bespoke tools in isolation, engineers can rely on a consistent framework to power Pinterest Assistant, ranking algorithms and safety systems. This shared baseline ensures that as Pinterest adopts increasingly sophisticated models that interpret text and images simultaneously, the entire ecosystem scales on a reliable footing.




