
Meta Collaborates with Cerebras to Drive Fast Inference for Developers in New Llama API
SUNNYVALE, Calif., May 2, 2025 — Meta has teamed up with Cerebras to offer ultra-fast inference in its new Llama API, bringing together the world’s most popular open-source models, Llama, with the world’s fastest inference technology, delivered by Cerebras.
Developers building on the Llama 4 Cerebras model in the API can expect generation speeds up to 18 times faster than traditional GPU-based solution. This acceleration unlocks an entirely new generation of applications that are impossible to build on other technology. Conversational low latency voice, interactive code generation, instant multi-step reasoning, and real-time agents — all of which require chaining multiple LLM calls — can now be completed in seconds rather than minutes.
By partnering with Meta to serve Llama models from Meta’s new API service, Cerebras gains exposure to an expanded global developer audience and deepens its business and partnership with Meta and their incredible teams.
Since launching its inference solutions in 2024, Cerebras has delivered the world’s fastest Llama inference, serving billions of tokens through its own AI infrastructure. The broad developer community now has direct access to a robust, OpenAI-class alternative for building intelligent, real-time systems — backed by Cerebras speed and scale.
“Cerebras is proud to make Llama API the fastest inference API in the world,” said Andrew Feldman, CEO and co-founder of Cerebras. “Developers building agentic and real-time apps need speed. With Cerebras on Llama API, they can build AI systems that are fundamentally out of reach for leading GPU-based inference clouds.”
Cerebras is the fastest AI inference solution as measured by third party benchmarking site Artificial Analysis, reaching over 2,600 token/s for Llama 4 Scout compared to ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.
Developers will be able to access to the fastest Llama 4 inference by selecting Cerebras from the model options within the Llama API. This streamlined experience will make it easy to prototype, build, and scale real-time AI applications. To sign up for early access to the Llama API and to experience Cerebras speed today, visit www.cerebras.ai/inference.
About Cerebras Systems
Cerebras Systems is a team of pioneering computer architects, computer scientists, deep learning researchers, and engineers of all types. We have come together to accelerate generative AI by building from the ground up a new class of AI supercomputer. Our flagship product, the CS-3 system, is powered by the world’s largest and fastest commercially available AI processor, our Wafer-Scale Engine-3. CS-3s are quickly and easily clustered together to make the largest AI supercomputers in the world, and make placing models on the supercomputers dead simple by avoiding the complexity of distributed computing. Cerebras Inference delivers breakthrough inference speeds, empowering customers to create cutting-edge AI applications. Leading corporations, research institutions, and governments use Cerebras solutions for the development of pathbreaking proprietary models, and to train open-source models with millions of downloads. Cerebras solutions are available through the Cerebras Cloud and on-premises.
Source: Cerebras
May 2, 2025
- Meta Collaborates with Cerebras to Drive Fast Inference for Developers in New Llama API
- Coalesce Honors Data Leaders Driving Innovation with 2025 GOAT Awards
- Boomi Expands DataHub with Unified Command Center for Cross-Platform Data Management
- LlamaIndex Announces Investments from Databricks and KPMG LLP
- BigID Launches End-to-End Data Lifecycle Management to Tackle AI and Compliance Risks
May 1, 2025
- Adastra Named AWS Data Foundation Partner, Helping Organizations Ready Their Data for GenAI
- Treasure Data Achieves Google Cloud Ready – BigQuery Designation
- Linux Foundation Expands AI Tooling with 3 IBM-Backed Open Source Projects
- GigaIO Partners with d-Matrix to Deliver Ultra-Efficient Scale-Up AI Inference Platform
- Precisely: New Global Research Reveals Key Observability Trends and Challenges for AI Innovation
- GridGain Sponsors Leading Data, Analytics and AI Industry Events
- Astronomer Secures $93M Series D Funding to Deliver Unified DataOps Platform for Enterprise AI
- Fivetran Signs Agreement to Acquire Census
- Akka Launches New Deployment Options for Agentic AI at Scale
April 30, 2025
- LogicMonitor Expands AI Observability Platform with Agentic AIOps and New Partnerships
- KNIME Turns Enterprise Data into Action, Demonstrates Custom AI Agents
- Pythian Boosts Global Data and AI Services with Rittman Mead Integration
- Collibra Harris Poll Finds 86% of Data Leaders Cite Privacy as Top Concern Amid AI Adoption
- StarTree Adds AI-Native MCP and Vector Embedding to Power Real-Time RAG and Agentic Apps
- DDN and Nebius Partner to Deliver Scalable AI Infrastructure for Enterprise Applications
- PayPal Feeds the DL Beast with Huge Vault of Fraud Data
- Thriving in the Second Wave of Big Data Modernization
- OpenTelemetry Is Too Complicated, VictoriaMetrics Says
- Google Cloud Preps for Agentic AI Era with ‘Ironwood’ TPU, New Models and Software
- Google Cloud Fleshes Out its Databases at Next 2025, with an Eye to AI
- Slash Your Cloud Bill with Deloitte’s Three Stages of FinOps
- Can We Learn to Live with AI Hallucinations?
- Monte Carlo Brings AI Agents Into the Data Observability Fold
- AI Today and Tomorrow Series #3: HPC and AI—When Worlds Converge/Collide
- The Active Data Architecture Era Is Here, Dresner Says
- More Features…
- Google Cloud Cranks Up the Analytics at Next 2025
- AI One Emerges from Stealth to “End the Data Lake Era”
- GigaOM Report Highlights Top Performers in Unstructured Data Management for 2025
- SnapLogic Connects the Dots Between Agents, APIs, and Work AI
- Supabase’s $200M Raise Signals Big Ambitions
- Snowflake Bolsters Support for Apache Iceberg Tables
- Dataminr Bets Big on Agentic AI for the Future of Real-Time Data Intelligence
- Big Data Career Notes April 2025
- GenAI Investments Accelerating, IDC and Gartner Say
- Dremio Speeds AI and BI Workloads with Spring Lakehouse Release
- More News In Brief…
- Gartner Predicts 40% of Generative AI Solutions Will Be Multimodal By 2027
- AMD Powers New Google Cloud C4D and H4D VMs with 5th Gen EPYC CPUs
- Opsera Raises $20M to Expand AI-Driven DevOps Platform
- GitLab Announces the General Availability of GitLab Duo with Amazon Q
- BigDATAwire Unveils 2025 People to Watch
- Dataminr Raises $100M to Accelerate Global Push for Real-Time AI Intelligence
- SAS Partners with Kansas State to Advance AI-Driven Water Management
- SoftServe Partners with Google Cloud to Accelerate Agentic AI and Data Initiatives
- DataChat Achieves Google Cloud Ready – BigQuery Designation
- Denodo Powers Data Fabric for LG U+ Network Division Modernization
- More This Just In…