Prime Intellect has launched Prime Inference, a cutting-edge platform designed for serving frontier open models. This OpenAI-compatible service utilizes NVIDIA Blackwell technology, marking a significant advancement in AI model deployment.
The platform’s GLM-5.3 deployment incorporates innovative technologies such as Dynamo, vLLM, and NVFP4 KV compression. This allows it to efficiently serve up to 66 sessions per prefill group, achieving an impressive rate of 101 tokens per second per user.
With Prime Inference, Prime Intellect aims to enhance the accessibility and performance of AI models, supporting developers and researchers in their endeavors to leverage advanced AI capabilities.
Compiled automatically by the Tech AI Newsdesk from public AI-news sources and summarised in our own words.