Meta is going modular with its next-generation MTIA chips, betting design flexibility beats early tape-out on a rigid architecture.
Meta will begin production of its next-generation MTIA (Meta Training and Inference Accelerator) chips in September, using a modular design intended to let the architecture evolve as AI workload characteristics shift.
Custom AI silicon has historically been a rigid bet: pick an architecture 18 months out, ship it, hope the workload it optimizes for is still relevant. Meta's modular approach breaks that trade-off — different components can be swapped without redesigning the whole chip.
For Meta, the point isn't matching Nvidia's peak flops. It's cutting inference cost for the specific workloads Meta actually runs at scale — feed ranking, recommendations, video understanding. At Meta's inference volume, even 20% cost reduction is material.
Nothing in the near term. Meta will continue buying Nvidia for training. But the long-arc pattern — hyperscalers building custom inference silicon while keeping Nvidia for training — is accelerating. Google, Amazon, and Microsoft all have similar programs.
Source: TechCrunch