AI news ·
Apple unveils Core AI framework for on-device large language models across its hardware lineup
Apple unveiled Core AI at WWDC 26 as the successor to Core ML, offering a unified on-device architecture for LLMs from 3B to 70B parameters with zero server costs. Developers can deploy AI across iPhone, iPad, Mac, and Vision Pro using a single Swift API with no per-token cloud dependencies.

Core AI is the engine beneath Apple Intelligence. With the next OS and toolchain releases, Apple is opening that engine to developers for what it calls "custom intelligence." The framework runs exclusively on Apple Silicon and enforces on-device data privacy by design.
What the framework gives developers
The API unifies hardware access so workloads move across CPU, GPU, and the Neural Engine through a single interface. A memory-safe Swift API provides zero-copy data paths and fine-grained control over inference memory. Ahead-of-time compilation shifts preparation work off the device, producing near-instant load times when a user launches a model.
Model compression applies optimization techniques like quantization and palettization, aligned by default with the Core AI runtime's execution patterns. Compression can reduce disk footprint, runtime memory, inference latency, and power consumption-or target all four at once.
Three paths for model deployment
Apple now supports three distinct approaches on its operating systems. According to developer discussions, the guidance places Core ML for classic, non-neural machine learning such as decision trees or tabular feature engineering. Core AI handles neural networks and transformers. MLX Swift is for working with custom model weights, though potentially with lower performance.
Community feedback notes that while Core AI "makes it easier to incorporate high-performance LLMs," its long-term value will depend "on the future growth of the official Core AI/community."
Converting and optimizing PyTorch models
Developers can bring a PyTorch model into Core AI by exporting it as a torch.export.ExportedProgram and converting it to a CoreAI AIProgram using TorchConverter().add_exported_program(ep).to_coreai(). For deeper customization, the framework offers built-in composite ops-attention, RoPE embeddings, RMSNorm, gather-matmul-along with the ability to register custom lowering functions that map new PyTorch ops to the Core AI intermediate representation, or write custom Metal kernels for low-level optimization.
One critical runtime behavior is automatic specialization. When a model first loads into the model cache, it specializes to the current hardware and OS version. That first run may take significantly longer than subsequent calls. Developers can manage this through SpecializationOptions, check or clear cached models via AICacheModel, and share the model cache across an app group.
Why this matters for product development
Core AI removes the server-side dependency that has defined most LLM integration to date. For product teams building features around Generative AI and LLM capabilities, zero per-token cloud costs and on-device privacy change the unit economics and the compliance surface. The trade-off is hardware lock-in: Core AI runs only on Apple Silicon. But for apps targeting Apple's ecosystem, the framework provides a single API that scales from a 3B vision model on an iPhone to a 70B reasoning model on a Mac, with compression and caching tooling built directly into the pipeline. Teams working on AI for IT & Development will need to evaluate whether Core AI's unified hardware model and AOT compilation offer enough performance advantage over MLX for their specific workloads.