编程 Core AI 替换 Core ML:.mlpackage → .aimodel 的转换流水线与 Swift 运行时

2026-10-10 00:04:05

Core AI 替换 Core ML:.mlpackage → .aimodel 的转换流水线与 Swift 运行时

Core AI 是 Apple 的 on-device 推理框架,随 WWDC 2026 / iOS、macOS 27 出现,定位是 Core ML 的后继者,也是设备上跑 Apple Intelligence 的推理框架,现在对开发者开放。它保留「convert once, run on ANE/GPU/CPU」的思路,但用新的 IR、compiler、runtime 整体替换旧的 .mlpackage + coremltools 栈:转换产物从 .mlpackage 变成 .aimodel,转换工具从 coremltools 换成 coreai-torch,优化从 coremltools.optimize 换成 coreai-opt。

组件构成

  • CoreAI.framework:设备端 Swift API,内存安全,负责加载 .aimodel 并在 Apple Silicon 的 CPU/GPU/Neural Engine 上执行推理。核心类型有 AIModel(加载 .aimodel)、InferenceFunction(可运行的计算图,通常是 main)、NDArray(多维输入输出数据)、NDArray.View / NDArray.MutableView(内存安全的数组访问)。
  • coreai-core:Python wheel,pip install coreai-core,用于构建和运行 AI Model。内含 coreai.authoring(从 Python 构建 AI Model)与 coreai.runtime(加载并执行 .aimodel,接受 NumPy 输入)。compiler 和 runtime 闭源,随 coreai-core wheel 以及 OS 里的 CoreAI.framework 一起发布。
  • coreai-torch:Core AI 的 PyTorch 扩展,把 PyTorch 模型转成 Core AI asset。pip install coreai-torch 会同时安装 coreai 包和 coreai-torch 库,对应旧栈里的 coremltools converter。
  • coreai-optimization(coreai-opt):量化、palettization、剪枝,基于 torchao PT2E,对应 coremltools.optimize。可以按层选择压缩技术、调粒度。
  • coreai-models:Apple 的模型 zoo + Swift runtime + agent skills,开源在 github.com/apple/coreai-models,提供可直接跑的 Swift package,在 App 里跑 LLM。
  • Core AI Debugger + Xcode 集成:Inspect Core AI graphs、profile、部署前验证 artifact。Core AI Debugger 是独立 macOS 应用,能把 tensor 值直接追溯到原始 Python 源码;运行时性能用 Instruments 分析。

PyTorch → .aimodel 三步流水线

  1. torch.export.export(model, args=(...)) 捕获计算图,得到 ExportedProgram(已针对示例输入做了 shape/dtype 特化)。
  2. ep.run_decompositions(get_decomp_table()):get_decomp_table() 返回默认 ATen 分解表,但扣掉了 TorchConverter 会当 composite op 降级的算子,所以这些算子会保留在导出图里,不被拆成底层原语。
  3. TorchConverter().add_exported_program(ep).to_coreai() 得到 AIProgram,再执行 coreai_program.optimize()。TorchConverter 逐节点遍历 FX graph,把 ATen 算子映射到 Core AI 算子。
import torch
from coreai_torch import TorchConverter, get_decomp_table

model = MyModel().eval()
ep = torch.export.export(model, args=(torch.randn(1, 10),))
ep = ep.run_decompositions(get_decomp_table())
coreai_program = TorchConverter().add_exported_program(ep).to_coreai()
coreai_program.optimize()

工作流按手上的输入选:

  • 已有分解好的 ExportedProgram → TorchConverter().add_exported_program(ep).to_coreai()
  • 有 nn.Module、不需要 externalization → add_exported_program 或 add_pytorch_module
  • 有 nn.Module、需要 externalization → add_pytorch_module(model, ..., externalize_modules=[...])

.aimodel 结构与 stateful 执行

.aimodel 是一个目录 bundle,内容为 {metadata.json, main.mlirb, main.hash}(IR + manifest)。它可以包含多个 function(entrypoint),并声明 states —— 图内会被原地修改的 tensor,运行时通过 state= API 暴露。KV cache 就是这么实现的:transformer 的 key/value cache 避免重算全部历史,降低随序列长度增长的推理延迟。

authoring 与自定义算子

coreai_torch.composite_ops 暴露了常用模块:attention、RoPE embeddings、RMSNorm、gather-matmul(MoE 原语)、SDPA、GatedDeltaUpdate 等。把模块传给 externalize_modules 会保留其算子边界为具名 composite op,让 compiler 能识别并优化。

  • 没有内置 lowering 规则的 torch op:用 register_torch_lowering 注册自定义 lowering。
  • 计算密集型自定义 op:用 register_custom_kernels / TorchMetalKernel 直接写 Metal 4 kernel 源码,接入转换流水线。
  • 高级用法:用 target-aware patterns/layouts 重新 author 模型,针对具体设备家族榨取能效。

specialization 与 AOT 编译

.aimodel 是 source 表示,能在 Apple 设备上跑,但必须先针对用户的具体设备做 specialization 才能加载推理。大模型首次 specialization 可能很耗时,之后从 cache 加载很快(具体耗时取决于模型规模和设备,官方未给出基准数据,此处未实测)。相关 API:AIModelCache(默认模型缓存)、显式请求 specialization、SpecializationOptions(配置行为)、删除无用条目、控制持久化,以及在同一 app group 内跨多个 App 共享 cache。

AOT 编译可以把部分编译工作挪到开发机上。编译后的模型仍需设备特化,但设备端剩余工作更少。

部署上的建议:

  • 用 cache 检查来 gate 功能,或提示用户 AI 功能正在准备中;
  • 在下载完 asset 或用户开启功能之后,再请求 specialization;
  • 不要在用户交互流程内做 specialization。

安装

pip install coreai-torch   # 含 coreai + coreai-torch
pip install coreai-core    # coreai.authoring / coreai.runtime,NumPy 输入

第三方观察

john-rocky/coreai-model-zoo 指出,Apple 的 coreai-models 模型 zoo 落后约一代(Qwen3 / Gemma 3,无 VLM),Swift runtime 也假设的是标准 input_ids → logits + 单 KV 的模型。更新的架构,比如 hybrid linear-attention SSM、dual-KV + per-layer-embedding decoder、VLM,需要重新 author,并写能处理非标准 state 的 runner。

复制全文 生成海报 core ai CoreML PyTorch Apple Silicon

推荐文章

程序员茄子在线接单