Skip to content
View in the app

A better way to browse. Learn more.

Armbian Community Forums

A full-screen app on your home screen with push notifications, badges and more.

To install this app on iOS and iPadOS
  1. Tap the Share icon in Safari
  2. Scroll the menu and tap Add to Home Screen.
  3. Tap Add in the top-right corner.
To install this app on Android
  1. Tap the 3-dot menu (⋮) in the top-right corner of the browser.
  2. Tap Add to Home screen or Install app.
  3. Confirm by tapping Install.

ODROID-M1: RK3568 NPU on the open stack (rocket kernel driver + Mesa Teflon)

Featured Replies

The NPU on the ODROID-M1 can execute ordinary quantized TFLite models through the standard TFLite delegate mechanism, on a fully open stack: the `accel/rocket` kernel driver (with local fixes) plus the Mesa Teflon delegate. An application loads a regular .tflite file — no model-conversion step, no proprietary runtime. Image classification (mobilenet v1/v2, resnet18) and object detection (yolov8n) work today, and other convolutional networks take the same path. This complements the vendor-stack setup from [the DKMS topic](https://forum.armbian.com/topic/58036-odroid-m1-npu-fully-working-on-armbian-618x/), on the same board.

What runs, measured against CPU-only TFLite on the same board:

- **mobilenet_v1** (224×224 quant): **5.4 ms vs 117 ms (~22×)** — convolutions, depthwise, the fully-connected final layer and average pooling all on the NPU; top-5 classes match the CPU reference, mean absolute logit difference 0.03, bit-identical results across repeated runs.
- **mobilenet_v2**: **6.3 ms vs 76 ms (~12×)** — including all ten fused residual additions; top-5 matches the CPU reference including order.
- **resnet18**: **11.2 ms vs 345 ms (~31×)** — end-to-end in a single NPU partition: convolutions, residual adds, max pooling, global average pooling and the fully-connected classifier, the poolings executed by the NPU's dedicated PPU unit.
- **yolov8n** (320×320 int8, the plain ultralytics TFLite export): **14.9 ms vs 70 ms (~4.7×)** — backbone, neck and the three detection heads on the NPU in two partitions, including all 57 SiLU activations (fused into the convolutions through the DPU lookup table, as the vendor stack does) and the two nearest-neighbour upsamples (the DPU's unpooling mode); the NPU itself takes 8 ms, the rest is the [1, 84, 2100] DFL tail on the CPU. On bus.jpg the same five detections as the CPU (boxes, classes, scores).
- Single-operation probe layers (regular, depthwise, stride-2, the 3-channel first layer, fully-connected, concat, standalone add, SiLU, slice/pad/requantize views, upsample — uint8 and int8, per-tensor and per-channel weights) match the TFLite CPU reference within 1 LSB.
- For the classifiers only reshape and softmax stay on the CPU.

Timings are at the NPU clock pinned to 800 MHz through SCMI (the MAC array is clocked from the SCMI/PVTPLL rate, not the CRU clock the driver sets; a module parameter selects the rate). "Bit-identical" means: same board, fixed clock, repeated runs; results from other RK3568 boards are welcome.

How the hardware is driven: the whole delegated graph is submitted as one job. Each task's command stream is patched to chain into the next, the NPU's PC unit walks the chain in hardware and raises a single interrupt at the end — the same submit model the vendor driver uses. Per-job overhead is ~0.2 ms, so avoiding a round-trip per layer matters: it is most of the difference between these numbers and a naive per-task submit. The NPU cannot address memory above 4 GiB, but no `mem=4G` boot restriction is needed: GFP_DMA32 patches in the IOMMU page-table path and in the driver's buffer allocation keep the full 8 GB usable.

This builds on the RK3588 Mesa branch (MR !42134). RK3568 differs: 8-channel feature atomics instead of 16 (strides in 8-byte units), 8×32 KiB CBUF banks, a different weight layout (with compact tails for partial channel slices and kernel groups), depthwise above 32 channels running as 32-channel-group tasks, a packed-RGB first-layer mode, weight streaming for FC-shaped layers, requantization through the DPU BS stream, a feature-prefetch (FEATURE_GRAINS) formula that grows on narrow feature maps, the element-wise unit configuration for residual adds — including the surface-notch addressing that band-split add tasks need — the PPU/PPU_RDMA pooling programming, the DPU lookup table (SiLU; the tables travel with the job and the kernel writes them over MMIO, loading them from the command stream lands only the first one) and the DPU unpooling mode (2× upsample), all derived from byte-level comparison against captured vendor command streams (resnet18 and probe models). Details are in the commit messages.

Code and reproduction (step-by-step README, kernel .debs in the release):
- https://github.com/iav/rk3568-npu — kernel module, DT overlay, test harness, probe models, README
- https://github.com/iav/mesa (branch rk3568-test-session-20260820)
- release with the kernel .debs / rocket.ko / dtbo: https://github.com/iav/rk3568-npu/releases/tag/v2026.08.22

References used: MR !42134 (https://gitlab.freedesktop.org/mesa/mesa/-/merge_requests/42134), the vendor-stack topic above (incl. the >4 GB IOMMU discussion), kiln (https://github.com/gahingwoo/kiln), rknn-toolkit2 (https://github.com/airockchip/rknn-toolkit2 — only needed for further RE, not for running).

Limitations: residual ±1-per-layer rounding drift (the hardware rounding pipeline differs from TFLite's reference requantization); PPU pooling covers kernels ≤ 8 with equal strides and equal input/output quantization, other average poolings fall back to a depthwise-convolution lowering; softmax and the 3D detection tail on the CPU; convolution/depthwise/add/pooling/concat/fully-connected with fused ReLU/ReLU6/SiLU, slice/pad/requantize as addressing, nearest 2× upsample; uint8/int8 activations (per-tensor), per-tensor or per-channel weights; feature maps up to 2047 wide/high. A rare (~2% of fresh processes) one-pixel race between chained tasks shows on tiny probe graphs only; the real networks are bit-identical run to run. The Mesa branch head is ahead of the patch series in the rk3568-npu repo; use the branch for the latest state.

The 13× MobileNet V1 number matches what we see on RK3588 at the same clock — good to see
it confirmed on RK3568 too.

Two small things worth checking before the mesa branch goes anywhere:

1. PC_TASK_CON.task_number is 8-bit on RK3568 vs 12-bit on RK3588. Make sure the rocket
driver masks it, or commands above 255 truncate silently.

2. If you're claiming bit-identical results, pin the DVFS point and CMA region first — the
same model can drift across boards otherwise, and this starts to matter more once
RK3576 DVFS lands upstream.

We have RK3568 boards here (SBC3568) and can run the same MobileNet V1/V2 + sha256 oracle
for comparison if useful. Do you plan to send the mesa branch upstream, or keep it downstream for now?

  • Author

Thanks. A run on SBC3568 would help — use release v2026.08.22 + README, we'll compare output sha256s.
task_number: vendor rknpu_drv.c sets pc_task_number_bits=12 for RK3568, same as RK3588, and the register confirms it — write 0xfff, read back 0xfff.
"Bit-identical" is at a clock pinned via SCMI (800 MHz) — added to the post above.
No upstream submission from my side — no resources to work with upstream. The branch is public, anyone is welcome to pick it up.

Hi,
 

Many thanks for your hard work on the Mesa patches and for adapting the stack to the RK3568 platform. It's great to see it running on the ODROID-M1. Keep it up! 🙂

I managed to get it working on my ODROID-M1 (8 GB) systems running my own Armbian build based on v26.08 and 7.2.0-rc7-bleedingedge-rockchip64 with this iommu patch:

 

(npu-test) fero@odroidtest:~/rk3568-npu$ sudo dmesg | grep -w npu
[    9.743716] rockchip-pm-domain fdd90000.power-management:power-controller: failed to get ack on domain 'npu', val=0x1ee
[   10.753390] platform fde40000.npu: Adding to iommu group 3
[   14.203903] rocket fde40000.npu: npu-supply enabled at 850000 uV
[   14.204531] rocket fde40000.npu: TEST scmi npu clk: enable=0 set_rate=0 rate=800000000
[   14.215487] rocket fde40000.npu: raw VERSION=0x00000000 VERSION_NUM=0x00000000
[   14.216241] rocket fde40000.npu: Rockchip NPU core 0 version: 0


and here are the results from some of your test scripts:
 

(npu-test) fero@odroidtest:~/rk3568-npu$ python tests/suite.py models/layer-c28.tflite
INFO: Created TensorFlow Lite XNNPACK delegate for CPU.
layer-c28.tflite f0:1/0.09 f128:1/0.22 f255:1/0.06 colgrad:1/0.06 rowgrad:1/0.07 chan:1/0.20 mix:1/0.06

(npu-test) fero@odroidtest:~/rk3568-npu$ python tests/mob.py
INFO: Created TensorFlow Lite XNNPACK delegate for CPU.
mobilenet_v1_1.0_224_quant.tflite cpu 111ms npu 5.4ms top5 5/5 logit diff max 16 mean 0.03 | npu top5 [412, 742, 736, 886, 912]


 

  • Author

@ferro, thanks for the run — 7/7 layer probes and v1 at 5.4 ms with top-5 5/5 on a different board and kernel build is exactly the confirmation that was missing. Your iommu patch is the same GFP_DMA32 fix baked into the released .debs (`iommu_data_ops_v2` page tables land above 4 GB on 8 GB boards without it); it now ships in the repo as `kernel-patches/` so nobody has to hunt for it.

Hi iav, Thanks — taking your offer.
We'll run release v2026.08.22 on a Boardcon SBC3568 (RK3568 SBC, 8 GB, Murata 1XD module for Wi-Fi, eMMC boot) tonight, using the README's standard path: install the kernel .debs, rocket.ko, and DT overlay, then patch in the GFP_DMA32 IOMMU fix from your `kernel-patches/` directory. We'll match your oracle — SHA-256 of the NPU output against the CPU reference for each of the seven layer-probe models plus mobilenet_v1 / v2 / resnet18, with the result captured at 800 MHz via SCMI (your `scmi_rate=800000000` module parameter) and DVFS pinned.
Specifically we want to cover three angles that ODROID-M1 doesn't quite match — eight-GiB usable memory above 4 GiB, the Murata radio coex path that shares an IOMMU group with the NPU on this board, and a long-running thermally-constrained soak at 70°C ambient rather than room temperature. The last one is what we do all the time here, so it falls out for free. Will post back in 48 h with: - the per-model SHA-256s and timings - top-5 match against the CPU reference - one-hour therm soak log (CPU/NPU temperature traces + NPU clock held at 800 MHz via SCMI despite thermal pressure — the same bit-identity check you ran at room temperature) If results hold, would you like us to add the report as a Tested-by line in your v2026.08.22 README, or are you sending that as a separate Mesa branch PR?

  • Author

That's more than I could test myself — the 70 °C soak and the radio coex sharing an IOMMU group with the NPU are exactly the two angles I have no hardware for; everything I measured was one board at room temperature.

I'd be glad to add your results to mine: a Tested-by line in the v2026.08.22 README plus your full report (SHA-256s and traces), linked back to this post.

No Mesa PR from me, though — I have no resources for upstreaming. The branch is public under the same licence as Mesa; if someone else wants to carry it upstream I'd be happy, and I'll answer questions on the details, but I won't be driving it.

Our BSP engineer kicked off the SBC3568 Armbian SDK build (legacy BSP 5.10 branch via compile.sh, first time on this board), and we're hitting the usual first-build friction — dependency resolution, WSL2 environment quirks, and a few Rockchip-BSP-specific toolchain issues. Nothing show-stopping, just first-time setup pain. Realistic timeline: 3-4 working days to get a clean .debs set, which puts our data delivery at ~Aug 31 – Sep 2. I'd rather flag this now than go silent past the 48h mark and show up late with no heads-up. The two test angles you called out — 70°C sustained soak and the radio coex / NPU IOMMU group interaction — are exactly what we're building toward. We're not cutting scope, just the build took longer than estimated. Will post incremental progress as the build advances.

  • Author

@Boardcon_yang, thanks for the heads-up on the SBC3568 build — the timeline is fine.

The branch matters more than the delay: `accel/rocket` is a mainline DRM/accel driver and does not exist in the Rockchip legacy BSP 5.10 tree, so even a clean BSP .debs set won't carry this stack. The README targets Armbian `bleedingedge` rockchip64 7.2-rc7, and the .debs in release v2026.08.22 are from that.

A full image build isn't needed either — `./compile.sh kernel BRANCH=edge` with your board config is enough, and it avoids most of the WSL2 friction: loop devices and binfmt only matter for the rootfs/image stages.

  • Author

@BipBip1981 — on the missing cpufreq_dt module: nothing broke. It was switched
from module to built-in in Armbian on 2026-02-28 (commit 6868e1c9, "orangepi5:
fix slow boot on current kernel"), so rockchip64 current/edge carry
CONFIG_CPUFREQ_DT=y. lsmod will never show it, and policy0/policy4 being present
means the driver is active. Nothing to restore.

More useful is what your own tests show: cpufreq.off=1 is what gets the board
through boot, with either overlay or none. That points at CPU DVFS/OPP, not
PCIe. The PERST# overlay did apply on your board (gpio-28 (|ep) out hi) and did
not change your crashes — so your failure class is not the boot SError I was
chasing.

  • Author

@amreo — no HDDs and no PCIe controller at all, both with and without the patch,
is a different story: the overlay only changes reset timing on a link that
otherwise trains. Could you post dmesg | grep -i pcie and your kernel version?

Join the conversation

You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.

Guest
Reply to this topic...

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.