Saturday, June 28, 2025

AI Inference: Meta Groups with Cerebras on Llama API


Sunnyvale, CA — Meta has teamed with Cerebras on AI inference in Meta’s new Llama API, combining  Meta’s open-source Llama fashions with inference know-how from Cerebras.

Builders constructing on the Llama 4 Cerebras mannequin within the API can anticipate speeds as much as 18 instances sooner than conventional GPU-based options, in accordance with Cerebras. “This acceleration unlocks a wholly new era of purposes which might be not possible to construct on different know-how. Conversational low latency voice, interactive code era, instantaneous multi-step reasoning, and real-time brokers — all of which require chaining a number of LLM calls — can now be accomplished in seconds slightly than minutes,” Cerebras stated.

By partnering with Meta to serve Llama fashions from Meta’s new API service, Cerebras positive aspects publicity to an expanded developer viewers and deepens its enterprise and partnership with Meta and their unbelievable groups.

Since launching its inference options in 2024, Cerebras has delivered the world’s quickest Llama inference, serving billions of tokens via its personal AI infrastructure. The broad developer group now has direct entry to a strong, OpenAI-class different for constructing clever, real-time methods — backed by Cerebras pace and scale.

“Cerebras is proud to make Llama API the quickest inference API on this planet,” stated Andrew Feldman, CEO and co-founder of Cerebras. “Builders constructing agentic and real-time apps want pace. With Cerebras on Llama API, they will construct AI methods which might be essentially out of attain for main GPU-based inference clouds.”

Cerebras is the quickest AI inference answer as measured by third social gathering benchmarking web site Synthetic Evaluation, reaching over 2,600 token/s for Llama 4 Scout in comparison with ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.

Builders will be capable to entry to the quickest Llama 4 inference by choosing Cerebras from the mannequin choices inside the Llama API. This streamlined expertise will make it simple to prototype, construct, and scale real-time AI purposes. To join early entry to the Llama API and to expertise Cerebras pace as we speak, go to www.cerebras.ai/inference.



Related Articles

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Latest Articles

PHP Code Snippets Powered By : XYZScripts.com