Npunlock Writes Custom C Kernels for Intel's NPU
Hardware / explainer
Npunlock Writes Custom C Kernels for Intel's NPU
The four-day-old GitHub project reconstructs a path Intel's own driver never exposes, running hand-written C on the NPU's programmable SHAVE cores.

What problem is npunlock solving?
Intel's Core Ultra processors ship a neural processing unit built around programmable SHAVE cores, the same streaming-hybrid architecture Intel inherited when it bought Movidius in 2016. Intel's own linux-npu-driver documentation describes a driver stack built on the oneAPI Level-Zero API, where a program calls zeGraphCreate to compile a graph of operations the driver already supports, and the compiled result is cached to disk. Nowhere in that stack is there a public workflow for supplying a custom, hand-written operation. If the operation you need is not one of the ones Intel's compiler already knows, the public NPU stack has no answer for you.
npunlock, a project published on GitHub on Sept. 20 by a developer using the handle hsfzxjy, closes that gap. Its own description states the problem it was built to solve: "Intel ships programmable SHAVE cores inside its NPUs, but the public stack exposes only graph-level programming." The project reconstructs, through reverse engineering rather than any Intel-published specification, a path from ordinary C source code to a compiled kernel that runs on the NPU's ACT-SHAVE processors, which are the same programmable cores the public stack refuses to expose directly.
How the workaround actually runs
The published example compiles a full-precision GELU activation function, the nonlinear operation used throughout transformer models, from 20 lines of C. That C source calls a small set of macros from npunlock's own header to read the NPU's input tensor and write its output, computes 0.5 * x * (1 + tanh(...)) on each element, and gets embedded into an NPU graph alongside operations that still come from Intel's own compiler. Compiling that C to ACT-SHAVE machine code requires a separate toolchain, Intel/Movidius MoviTools' MVC_DEPEND component, which npunlock does not redistribute; its documentation instead points developers to extract the tool from a Lenovo NPU driver package numbered 31.0.100.1688, while warning them explicitly not to install that legacy driver itself. Once compiled, the kernel runs on real NPU hardware and its output is checked against a NumPy reference computed on the CPU.
The project's most recent capability, added Sept. 23, lets one graph run two custom branches at different precisions at once, an FP32 unary path alongside an FP16 binary path, which required working around how Intel's own compiler reorders ACT groups during compilation. npunlock is released under the Apache License 2.0; the MoviTools and Movidius components it depends on remain Intel's proprietary, unredistributed property.
What it can do today, and what it can't
Everything in the project has been verified on one configuration: Windows x64, running on Meteor Lake silicon carrying Intel's NPU3720. Intel's own driver, by contrast, lists support for five NPU-equipped processor generations. The gap between those two numbers is the honest scope of what has actually been tested.
| npunlock verified support | Intel's own linux-npu-driver | |
|---|---|---|
| Processor generations | 1 (Meteor Lake / NPU3720) | 5 (Meteor Lake, Arrow Lake, Lunar Lake, Panther Lake, Wildcat Lake) |
| Operating system | Windows x64 only | Linux |
| Custom C kernels | Yes, FP16 and FP32 unary/binary, static shapes | No public workflow |
The project's own documentation lists the same limits without softening them: static tensor shapes only, a fixed set of compatible "ACT carriers," and connected mixed-precision conversion groups that are "not yet patch-discoverable." A Linux port exists only as an unverified hypothesis in the project's wiki, built on the idea that a graph patched on Windows might run under Linux's NPU firmware unchanged, and on tooling the author calls "ready for hardware testing" but has not yet run against real Linux NPU passthrough.
Who built it, and how new this is
As of Sept. 24, the repository, hsfzxjy/npunlock, carries 50 stars, zero forks and zero open issues, according to GitHub's own API, four days after its first commit. That is early enough that adoption is still a question rather than an answer; it belongs in the same category as the open-source projects this site has tracked jumping GitHub's trending list on a given day, where a star count measures attention, not use. Intel's own driver documentation, for its part, describes no public plan to expose a custom-kernel path of its own, which is the specific gap that makes a single unaffiliated developer's four-day-old repository worth reading in the first place.
The project's README is direct about how it was built, in a section it labels a warning rather than a footnote: "I did use AI while building this project, for scaffolding, repetitive implementation work, converting my reverse-engineered results into organized documentation, and fixing my English. The reverse engineering, experiments, debugging, and technical conclusions came from hands-on work." That is a narrower and more specific disclosure than most hardware projects publish about their own tooling, including some Qualcomm has acquired outright without saying as much about how any of it was built.

What would make this matter beyond one laptop chip
The open question is whether the same reverse-engineered path generalizes past NPU3720. If a patched graph built on Windows runs unmodified on newer silicon because the SHAVE machine code itself is unchanged across generations, npunlock becomes a real alternative to Intel's graph-only stack across the five generations Intel already supports officially. If it doesn't, because newer NPUs use a different SHAVE image or incompatible firmware, this stays a Meteor Lake curiosity with an unusually well-documented reverse-engineering trail. Either way, that answer depends on hardware nobody has reported testing yet.
Sources
More in Hardware
- 01Tesla's Emergency Drive Away Lets Cars Leave a Supercharger Still Plugged In, Two Months After Twin FallsSoftware update 2026.38.3 lets drivers shift into Drive with the cable latched, at the cost of damage to car and charger that Tesla has not priced.
- 02GM's Equinox EV Fell More Than 90% in Q3, and Two Trackers Disagree on the Unit CountElectrek counts 1,905 Equinox EVs and GM Authority counts 1,705; the two also differ by 3,938 on the Blazer EV, which changes how large Cadillac's share looks.
- 03China Has Flown 70 Orbital Launches in 2026 and Needs 31 in Q4 to Pass 100Q3 closed at 26 launches, matching Q2; reaching 101 means a fourth quarter about 19% busier than either, with three commercial debuts and a crewed flight on the list.
- 04Atlas Gets a 13-Degree-of-Freedom Hand With No Pinky and No Published Grip ForceBoston Dynamics nearly doubled the hand's joints from 7 to 13 and aims at mass manufacture, but payload, torque, price and durability are all missing.