
Meta Collaborates with Cerebras to Drive Fast Inference for Developers in New Llama API
SUNNYVALE, Calif., May 2, 2025 — Meta has teamed up with Cerebras to offer ultra-fast inference in its new Llama API, bringing together the world’s most popular open-source models, Llama, with the world’s fastest inference technology, delivered by Cerebras.
Developers building on the Llama 4 Cerebras model in the API can expect generation speeds up to 18 times faster than traditional GPU-based solution. This acceleration unlocks an entirely new generation of applications that are impossible to build on other technology. Conversational low latency voice, interactive code generation, instant multi-step reasoning, and real-time agents — all of which require chaining multiple LLM calls — can now be completed in seconds rather than minutes.
By partnering with Meta to serve Llama models from Meta’s new API service, Cerebras gains exposure to an expanded global developer audience and deepens its business and partnership with Meta and their incredible teams.
Since launching its inference solutions in 2024, Cerebras has delivered the world’s fastest Llama inference, serving billions of tokens through its own AI infrastructure. The broad developer community now has direct access to a robust, OpenAI-class alternative for building intelligent, real-time systems — backed by Cerebras speed and scale.
“Cerebras is proud to make Llama API the fastest inference API in the world,” said Andrew Feldman, CEO and co-founder of Cerebras. “Developers building agentic and real-time apps need speed. With Cerebras on Llama API, they can build AI systems that are fundamentally out of reach for leading GPU-based inference clouds.”
Cerebras is the fastest AI inference solution as measured by third party benchmarking site Artificial Analysis, reaching over 2,600 token/s for Llama 4 Scout compared to ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.
Developers will be able to access to the fastest Llama 4 inference by selecting Cerebras from the model options within the Llama API. This streamlined experience will make it easy to prototype, build, and scale real-time AI applications. To sign up for early access to the Llama API and to experience Cerebras speed today, visit www.cerebras.ai/inference.
About Cerebras Systems
Cerebras Systems is a team of pioneering computer architects, computer scientists, deep learning researchers, and engineers of all types. We have come together to accelerate generative AI by building from the ground up a new class of AI supercomputer. Our flagship product, the CS-3 system, is powered by the world’s largest and fastest commercially available AI processor, our Wafer-Scale Engine-3. CS-3s are quickly and easily clustered together to make the largest AI supercomputers in the world, and make placing models on the supercomputers dead simple by avoiding the complexity of distributed computing. Cerebras Inference delivers breakthrough inference speeds, empowering customers to create cutting-edge AI applications. Leading corporations, research institutions, and governments use Cerebras solutions for the development of pathbreaking proprietary models, and to train open-source models with millions of downloads. Cerebras solutions are available through the Cerebras Cloud and on-premises.
Source: Cerebras
August 12, 2025
- SAS: Questing for the Quantum AI Advantage
- Atlan Launches App Framework with 21 Leading Partners to Make Context Shareable Across the AI-Native Enterprise
- Zilliz Expands Security and Compliance to Accelerate AI Adoption in Regulated Industries
- Sumo Logic Opens Submissions for 2nd Annual Sumie Awards
- Hitachi Vantara Recognized by GigaOm, Adds S3 Table Functionality to Virtual Storage Platform One Object
- Arcitecta Adds Native Vector Search and Metadata Integration to Mediaflux Platform
August 11, 2025
- HPE Helps Enterprises Drive Agentic and Physical AI Innovation With Systems Accelerated by NVIDIA Blackwell and the Latest NVIDIA AI Models
- NVIDIA RTX PRO Servers With Blackwell Coming to World’s Most Popular Enterprise Systems
- StorONE’s Efficient Platform Reduces Storage Guardian Data Center Footprint by 80%
- Dell Unveils Updates to Dell AI Data Platform
August 8, 2025
- Blaize Introduces AI Platform to Power Multi-Modal Intelligence at the Edge
- NCSA and Illinois Awarded $25.8M NGA Contract for HPC, AI, and Geospatial Data
- Quantiphi Achieves Google Cloud Data Management Specialization
- Computing Community Consortium Outlines Roadmap for Long-Term AI Research
August 7, 2025
- Oracle Helps Customers Achieve Extreme Availability and Performance for Mission-Critical and Agentic AI Applications
- Krutrim Partners with Cloudera to Power AI-Driven Innovation in India
- Elastic Introduces Logs Essentials: Serverless Log Analytics, in a New Low-priced Tier
August 6, 2025
- Scaling the Knowledge Graph Behind Wikipedia
- Rethinking Risk: The Role of Selective Retrieval in Data Lake Strategies
- Top 10 Big Data Technologies to Watch in the Second Half of 2025
- LinkedIn Introduces Northguard, Its Replacement for Kafka
- What Are Reasoning Models and Why You Should Care
- Apache Sedona: Putting the ‘Where’ In Big Data
- Top-Down or Bottom-Up Data Model Design: Which is Best?
- LakeFS Nabs $20M to Build ‘Git for Big Data’
- Doing More With Your Existing Kafka
- Why OpenAI’s New Open Weight Models Are a Big Deal
- More Features…
- Mathematica Helps Crack Zodiac Killer’s Code
- Supabase’s $200M Raise Signals Big Ambitions
- Promethium Wants to Make Self Service Data Work at AI Scale
- BigDATAwire Exclusive Interview: DataPelago CEO on Launching the Spark Accelerator
- Solidigm Celebrates World’s Largest SSD with ‘122 Day’
- The Top Five Data Labeling Firms According to Everest Group
- McKinsey Dishes the Goods on Latest Tech Trends
- Toloka Expands Data Labeling Service
- How AI Is Impacting the Job Market for College Grads
- AI Skills Are in High Demand, But AI Education Is Not Keeping Up
- More News In Brief…
- Seagate Unveils IronWolf Pro 24TB Hard Drive for SMBs and Enterprises
- Promethium Introduces 1st Agentic Platform Purpose-Built to Deliver Self-Service Data at AI Scale
- OpenText Launches Cloud Editions 25.3 with AI, Cloud, and Cybersecurity Enhancements
- TigerGraph Secures Strategic Investment to Advance Enterprise AI and Graph Analytics
- Gartner Predicts 40% of Generative AI Solutions Will Be Multimodal By 2027
- StarTree Adds Real-Time Iceberg Support for AI and Customer Apps
- Gathr.ai Unveils Data Warehouse Intelligence
- Databricks Announces Data Intelligence Platform for Communications
- Data Squared Announces Strategic Partnership with Neo4j to Accelerate AI-Powered Insights for Government Customers
- LF AI & Data Foundation Hosts Vortex Project to Power High Performance Data Access for AI and Analytics
- More This Just In…