Apple's Neural Engine Gets Its First Public Register Map
Hardware / analysis
Apple's Neural Engine Gets Its First Public Register Map
Independent researcher Eileen Yoon's account of the M1's accelerator shows Apple's 2020 marketing number holds only when the workload never leaves on-chip cache.
Independent researcher Eileen Yoon published a full reverse-engineered account of Apple's Neural Engine on Aug. 10, documenting the accelerator's compute datapath, scheduler, memory system and register-level control for the first time in the nine years since it debuted in the A11 Bionic chip.
The write-up covers the compiler output format CoreML emits to the chip, a fixed-size binary control packet Yoon calls the Task Descriptor, and firmware routines recovered through direct memory dumping. Yoon had already built an open-source ANE driver, posted on GitHub, but development on it stalled roughly three years before this account of the underlying hardware went up.
Nine years of TOPS claims, no register map until now
Apple has announced a headline throughput number for every generation of the chip without ever publishing what sits underneath it. The first one arrived Sept. 12, 2017, when Apple said the A11 Bionic's dual-core Neural Engine performed "up to 600 billion operations per second." By May 7, 2024, Apple said the M4's Neural Engine hit 38 trillion operations per second, which it called "60x faster than the first Neural Engine in A11 Bionic."
| Chip | Year | Neural Engine | Apple's claimed peak |
|---|---|---|---|
| A11 Bionic | 2017 | Dual-core | 600 billion ops/sec |
| M1 | 2020 | 16-core | 11 trillion ops/sec |
| M4 | 2024 | 16-core | 38 trillion ops/sec |
Yoon's reverse engineering only covers the M1 generation in the middle of that table. Six generations preceded it with no public register map at all, and every generation since, including the M4's 38 TOP/s design, still has none. Apple's developer documentation for CoreML has never described a Neural Engine's DMA bandwidth, cache size or register interface for any of them.
11 trillion operations per second, if the data never leaves cache
Apple's own M1 announcement on Nov. 10, 2020 said the 11 trillion operations per second gave the chip "up to 15x faster machine learning performance." Yoon's independent measurements land on that same 11 TOP/s peak. What Apple never published is what the peak requires: a sustained 68 GB/s of DRAM traffic, a ratio Yoon calculates at 162 operations per byte moved.
The chip cannot supply that on its own terms. Yoon measured the Neural Engine's two internal DMA engines directly: KernelDMA moves data at 37.99 GB/s and TileDMA at 59.08 GB/s, both below the 68 GB/s the peak figure assumes. For comparison, Yoon measured the M1's GPU moving data at 77.70 GB/s over the same memory system. A workload that keeps missing the Neural Engine's small on-chip caches will never touch 11 TOP/s. Apple's own developer documentation for CoreML does not describe any of these three numbers.
- Required for 11 TOP/s peak68 GB/s
- KernelDMA (measured)37.99 GB/s
- TileDMA (measured)59.08 GB/s
- M1 GPU, for reference (measured)77.7 GB/s
Source: Eileen Yoon, "Retrospectively Reverse-Engineering Apple's Neural Engine," Aug. 10, 2026
The memory hierarchy Apple has never documented
Yoon's account puts a size on every layer of the chip's cache. Each of the Neural Engine's 16 cores carries 64 KiB of local memory Yoon calls KMem, for 1 MiB total, backed by a 2 MiB cache shared across all 16 cores. Each core runs 128 parallel multiply-accumulate lanes at FP16 precision, or 256 at INT8, for 2,048 lanes across the whole chip. None of those figures, or the register addresses that configure them, appear in any Apple-published document; Yoon recovered them by dumping firmware memory and reading back the compiler's own hardware register files.
The scale of that gap matters for anyone building on-device inference today. A model whose working set fits inside roughly 2 MiB can plausibly approach the advertised throughput; the Qwen3.8-27B memory math The Terminal ran in a separate piece shows how quickly a modern context window blows past caches an order of magnitude larger than that, and the same arithmetic applies to any accelerator, not just Apple's.
A second, independent account, seven weeks earlier
Yoon's write-up is not the only one. Spencer Bryngelson, an assistant professor in Georgia Tech's School of Computational Science and Engineering, posted a 302-page account of the Neural Engine's architecture, programming model and performance to arXiv on June 23, seven weeks before Yoon's post went up. The two researchers worked independently, at different institutions, months apart, and arrived at the same conclusion neither Apple nor CoreML's documentation states outright: this is a real, measurable, memory-bandwidth-bound accelerator, not the black box its marketing treats it as. Bryngelson's paper runs to 302 pages precisely because so little of what it documents had ever been written down anywhere else, in any form, for any of the six chip generations that shipped before it.
Neither account required Apple's cooperation, and Apple has not published a response to either one. Both had to reconstruct behavior other accelerator vendors, including the ones behind browser-based inference hardware, typically document in a programming guide from day one. Nine years and six chip generations after the Neural Engine first shipped, the register map for the version in a 2020 laptop chip is only now public, and Apple has given no indication that documentation for the newer versions is coming.
Sources
More in Hardware
- 01Waymo Targets Singapore for 2028, Two Rivals Already Carry RidersWeRide and Pony AI have carried invited and paying riders through Singapore's Punggol district since April, roughly two years before Waymo's own timeline puts a rider in one of its cars there.
- 02Royal Enfield Prices Flying Flea at €5,990 Abroad, ₹2.79 Lakh at HomeNew Atlas pegs the electric motorcycle's April price in India at roughly $3,000 by direct conversion, and Royal Enfield has already lived through the same gap once with a gasoline model.
- 03Nvidia Won't Call Its Working Rust GPU Track Production-Readycutile-rs already backs an open-source LLM server and a Hugging Face testbed, but Nvidia's Sept. 8 announcement stops short of endorsing either new track for production.
- 04Arm Reuses the Total Design Name for Robots, Not Yet the SiliconThe original Total Design already has a customer-ready chiplet on TSMC's N2 process; the physical AI version Arm announced Sept. 8 is a set of robot-capability definitions.