Retrospectively Reverse-Engineering Apple's Neural Engine

Posted by zdw 8 hours ago

Counter181Comment19OpenOriginal

Comments

Comment by throw0101a 3 minutes ago

A lot of folks are/were saying that Apple has missed the boat when it comes to AI, in some (important) aspects that is correct, but I think it's worth remembering that Apple added the Neural Engine to A-series chips in 2017, before the AI hoopla really kicked off:

* https://en.wikipedia.org/wiki/Neural_Engine

* https://apple.fandom.com/wiki/Neural_Engine

"AI" has grown much more since then, and there have been important developments that Apple has not deployed (well), but I think they were looking ahead a little more than most at the time (even if events 'got away' from them subsequently).

Comment by zozbot234 7 hours ago

How does this relate to the more recent work on the M4 ANE found at https://maderix.github.io/articles/ ? Does the M4 and later ANE expose any additional capabilities, or is it just a higher-performance iteration of the same thing?

As an aside, the introduction to this article seems to conflate the ANE with the Neural Accelerators (NAX) found in the M5+ (and A-series equivalents) GPUs. These are very different things, and Apple is still working on the ANE - the M6 and A20 will apparently feature doubled ANE blocks.

Comment by woadwarrior01 6 hours ago

This one's authored by a human, the other one is authored by Claude.

> Does the M4 and later ANE expose any additional capabilities, or is it just a higher-performance iteration of the same thing?

IIUC, M4 introduced a fast path for INT8 weights and activations (w8a8). M5 Ultra, M6 and A20 have two ANEs.

> As an aside, the introduction to this article seems to conflate the ANE with the Neural Accelerators (NAX)

Yeah, that part is true. NAX cores are matmult accelerators, closer to tensor cores in NVIDIA GPUs.

Comment by anentropic 3 hours ago

Amazing analysis

Same author even found a bug in it https://eiln.github.io/posts/ane-dma.html

Comment by GeekyBear 3 hours ago

It's worth remembering that Apple is releasing a new framework (Core AI) this fall that goes beyond the Pytorch and Tensorflow workloads that the decade old Core ML framework allowed.

> Core AI allows your app to use the latest model architectures and inference techniques across the CPU, GPU, and Neural Engine.

https://developer.apple.com/documentation/coreai

Comment by CraigJPerry 8 hours ago

This isn't ai slop. It's fascinating and well written.

But I learned something really basic - i didn't know that the ANE (and the data pipeline around it) was designed for CNN rather than transformers. It's always been an open loop in my head, wondering why the ANE was less impactful than i understood it should be.

Comment by jasode 5 hours ago

>i didn't know that the ANE (and the data pipeline around it) was designed for CNN rather than transformers.

Multiple stories have reported that ANE came from Apple's self-driving car project that got canceled. (Makes sense since CNN is used for vision-related machine learning and enables cars to analyze their surroundings.) They spent 10 years and ~10 billion on research & development on a product that never got released so Apple is probably happy they're able to salvage some of that ai technology and put it in iPhones and Macs.

Comment by ACCount37 4 hours ago

Tesla also has its own NPUs for self-driving - and Tesla uses transformers for sensor fusion.

My guess would be that the main use case for an NPU in iPhone just used to be image processing/computational photography. Thus the CNN bent.

Also makes sense with the timing - back when iPhone first got its NPU, CV was the killer app for ML.

Comment by stefan_ 3 hours ago

This is pure sunk cost fallacy. CNNs were from the deep learning ImageNet heydays, but now everything in that domain is equally done better by transformers. All you are doing is wasting area and saddling software with outdated hardware, and myopic PMs insisting on its use will create inferior products. Now that sounds a lot like the Apple AI efforts..

Comment by riedel 7 hours ago

A lot of neural engine, particularly in the embedded domain (ARM/RISC MCUs) have the same problem. Designing other models means on top of this means a lot of profiling to get convolution blocks right to get good speedups. (We optimized this in the past e.g. using Neural Architecture Search on super networks)

Comment by msdz 7 hours ago

> But I learned something really basic

Same for me!

Also, just imagine being the group at Apple responsible for designing this section of the chip, starting probably almost a decade back – under the constant uncertainty of not knowing what direction ML workloads would develop in…

Comment by troupo 7 hours ago

ML research was a rather known quantity, or the separate "Neural Engine" CPU explicitly aimed at existing ML pipelines wouldn't exist.

However, very few used it for anything, even within Apple. I feel like it was a huge wasted opportunity.

Comment by eastbound 7 hours ago

I'm all for compassion, but engineers knew the NE was empty when it sat idle for 10 years on our computers.

- when you're given no usecase for your engineering piece, apart from "detour characters in pictures". It's an exageration but AI's contributions in iOS aren't visible; Meanwhile Google has features that people actually notice like removing tourists from your holidays photos — worse: it's mostly a simple collage feature working on the main CPU, and it has the same social effect as green bubbles in iMessage ("ah. Tourists on your photos. iPhone user?")

- and you tout it as "16 Neural Engine cores" during the sales, with no associated software, no listed material feature, just hand-waving,

- Siri maxxes out at "There is no contact named 'What's the weather today' in your agenda",

Then can't really claim that Apple engineers' problem was really the bad luck that ML wasn't the determining part of the future. It's more like misreading the room for 5 to 10 years straight.

Apple engineering's excellence on vertical integration and supply chain control gave them absolute power over our world (with merit), it just failed at that particular project. Which occupies 40% of our CPUs.

Comment by adastra22 5 hours ago

The ANE hasn't been sitting empty for 10 years. All those Photos features like face recognition and auto classification run on ANE.

Comment by hn9zmdcaou 6 hours ago

Ported a transformer to ANE and the whole job was pretending it was a CNN, 4D tensors with seq in the last axis and 1x1 convs instead of matmuls.

Comment by LoganDark 8 hours ago

> what workloads it was designed for and accels at.

excels!

Comment by pbhjpbhj 8 hours ago

Could've been a pun as a neural processor is an accelerator, so it 'accels' at machine learning tasks!

Comment by LoganDark 4 hours ago

Maybe, but then I wouldn't expect "at" :)

Comment by 4 hours ago

Comment by osquar 8 hours ago

At least we know it wasn't written by a bot

Comment by rima_667 7 hours ago

[flagged]

Comment by marbleotter115 6 hours ago

[dead]