According to The Information, Apple's in-house AI server chip, codenamed Baltra, has been delayed. The chip was originally planned for shipment in 2026 to support Apple's Private Cloud Compute infrastructure.
Behind the delay lies a severe performance bottleneck in existing infrastructure. Apple's current AI servers rely on the M2 Ultra chip used in Macs, which is highly efficient in personal computers but struggles with complex large-scale generative AI models.
The performance gap became evident when Apple engineers attempted to run Google's Gemini model locally on its own servers to revamp Siri—the hardware could not meet the required performance levels for the project.