Skip to content
View in the app

A better way to browse. Learn more.

Armbian Community Forums

A full-screen app on your home screen with push notifications, badges and more.

To install this app on iOS and iPadOS
  1. Tap the Share icon in Safari
  2. Scroll the menu and tap Add to Home Screen.
  3. Tap Add in the top-right corner.
To install this app on Android
  1. Tap the 3-dot menu (⋮) in the top-right corner of the browser.
  2. Tap Add to Home screen or Install app.
  3. Confirm by tapping Install.

[Armbian newsletter] - Mali GPU with LiteRT-LM

Featured Replies

Running Gemma 4 E2B on a Mali GPU with LiteRT-LM

Mali GPU with LiteRT-LM

This guide should work on any armbian minimal/console (trixie) system with a working Mali-g610 GPU.

DO NOT USE DESKTOP

Verified on: Orange Pi 5 Max (Rockchip RK3588, Arm Mali-G610 MC4), Armbian/Debian 13 (trixie), aarch64, 8 GB RAM.

Status summary: Both CPU (XNNPACK) and GPU (WebGPU/Dawn → Vulkan → panvk → panthor) work on the mainline panthor kernel — numbers below. The GPU path needed one line of local patching in
Mesa's panvk driver (raise maxImageDimension3D 512 → 2048 to meet the WebGPU minimums that LiteRT's Dawn library enforces; recipe in §5, rationale in §7.1). The model is multimodal (Text + Vision + Audio): the vision encoder runs on the same patched panvk via--vision-backend gpu (§5 "VL path" results). The vendor Rockchip libmali route and OpenCL remain unavailable on mainline panthor (§7.2, §7.3).

Stack (as shipped):

Gemma 4 E2B (.litertlm, int4)  →  LiteRT-LM CLI  →  LiteRT GPU accelerator (WebGPU/ML Drift)

Dawn (WebGPU impl) → Vulkan → panvk → panthor (kernel) → Mali-G610

1. Prerequisites

  • ARM (aarch64) Linux with a Mali GPU and at least 8 GB RAM.
  • Armbian/Debian 13 (trixie)
  • The panthor (or panfrost for older Mali) kernel driver active:
ls /dev/dri/renderD128          # render node must exist
lsmod | grep panthor            # or panfrost / lima

Full ARMv8.2-A dotprod support is required — LiteRT's ARM64 binaries are built with it and will SIGILL otherwise:lscpu | grep -w asimddp (both A55 and A76 cores list it here).

2. Install Mesa + Vulkan for Mali (needs root)

sudo apt update
# Prefer the newest Mesa (backports on Debian) for the best panvk coverage of Valhall CSF GPUs:
sudo apt install -y -t trixie-backports mesa-vulkan-drivers libvulkan1 vulkan-tools
# Optional (see §7 — currently yields NO usable OpenCL device on panthor):
sudo apt install -y -t trixie-backports mesa-opencl-icd clinfo

Verify the GPU is visible to Vulkan:

vulkaninfo --summary | grep -iE "deviceName|driverName"
# Expect: deviceName = Mali-G610 MC4   driverName = panvk

3. Install the LiteRT-LM CLI (user space, no root)

curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
uv tool install litert-lm        # current: v0.17.x; incl. CPU (XNNPACK/YNNPACK) + GPU (WebGPU/Vulkan) backends

4. Get the model (Gemma 4 E2B, int4 .litertlm, 2.58 GB)

mkdir -p ~/litert-lm-bench && cd ~/litert-lm-bench
curl -L -o gemma-4-E2B-it.litertlm \
  https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm/resolve/main/gemma-4-E2B-it.litertlm
# or have the CLI fetch it for you:
#   litert-lm benchmark --from-huggingface-repo litert-community/gemma-4-E2B-it-litert-lm gemma-4-E2B-it.litertlm

Other ready models: google/gemma-3n-E2B-it-litert-lm (gemma-3n-E2B-it-int4.litertlm),litert-community/gemma-4-E4B-it-litert-lm, ... (see litert-lm list / HF).

5. Benchmark

Lock the CPU governor to performance first (fair, repeatable numbers):

sudo sh -c 'for c in /sys/devices/system/cpu/cpu[0-7]/cpufreq/scaling_governor; do echo performance > $c; done'

CPU baseline (XNNPACK, 8 threads):

litert-lm benchmark ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
  -p 1024 -d 256 --backend cpu --cpu-thread-count 8 --cache disk --runs 2

GPU (patched panvk, §7.1). One-time local Mesa build — stays entirely in userland, the system Mesa is untouched:

# 1) Mesa source + one-line patch (26.1.2)
cd ~/litert-lm-bench
curl -L -o mesa.tar.xz https://archive.mesa3d.org/mesa-26.1.2.tar.xz
tar -xf mesa.tar.xz && cd mesa-26.1.2
python3 - <<'PY'
p = 'src/panfrost/vulkan/panvk_vX_physical_device.c'
s = open(p).read()
s = s.replace('.maxImageDimension3D = PAN_ARCH <= 10 ? (1 << 9) : (1 << 14),',
              '.maxImageDimension3D = (1 << 11), /* 2048 — meets WebGPU min */')
open(p, 'w').write(s)
PY

# 2) build deps (the LLVM/CLC chain is required — panvk bakes CLC-compiled SPIR-V for libpan)
sudo apt install -y --no-install-recommends meson ninja-build pkg-config gcc g++ python3-mako \
  python3-yaml python3-ply libdrm-dev llvm-19-dev libllvmspirvlib-19-dev spirv-tools \
  libclang-19-dev libclang-cpp19-dev

# 3) panvk-only build (~10–25 min on the 5 Max)
meson setup build-panvk -Dvulkan-drivers=panfrost -Dgallium-drivers= -Dbuildtype=release \
  -Dllvm=enabled -Dmesa-clc=auto -Dplatforms= -Degl=disabled -Dgbm=disabled -Dglx=disabled \
  -Dopengl=false -Dglvnd=disabled -Dtools= -Dbuild-tests=false
ninja -C build-panvk -j6

# 4) benchmark with the patched driver, process-scoped. First put the WebGPU prebuilts
#    (libLiteRtTopKWebGpuSampler.so + libwebgpu_dawn.so etc., litert-lm repo tag v0.17.1,
#    prebuilt/linux_arm64) in ~/litert-lm-bench/prebuilt_v0171/ — see Troubleshooting.
export BUILD=~/litert-lm-bench/mesa-26.1.2/build-panvk
export VK_ICD_FILENAMES="$BUILD/src/panfrost/vulkan/panfrost_devenv_icd.aarch64.json"
export LD_LIBRARY_PATH="$BUILD/src/panfrost/vulkan:$HOME/litert-lm-bench/prebuilt_v0171"
export PATH="$HOME/.local/bin:$PATH"            # uv-installed CLI
litert-lm benchmark ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
  -p 1024 -d 256 --backend gpu --cache disk --runs 2

First GPU run compiles the WebGPU/Vulkan shaders (one-time) and uploads GPU-rearranged weights (~0.8 GB weight cache); --cache disk persists both next to the model, so later loads start fast — the compile cost is NOT repaid on every run.

Results on this Orange Pi 5 Max (Mali-G610 MC4)

Backend Prefill (tk/s) Decode (tk/s) TTFT (s)
CPU (XNNPACK, 8 threads) 123.6 12.3 8.4
GPU (WebGPU/Dawn→Vulkan, patched panvk §5/§7.1) 374.0 8.7 2.9

Google's on-device-class references for Gemma 4 E2B: S26 Ultra GPU 3808/52 tk/s, Raspberry Pi 5 CPU 133/7.6 tk/s. On this board the GPU wins the prefill race by ~3.0× (374 vs 124 tk/s) and TTFT by ~2.9× (2.9 vs 8.4 s), while decode stays on the CPU's side (8.7 vs 12.3 tk/s) — as expected, decode is the bottleneck on Mali-class hardware; the Mali GPU is a prefill/TTFT accelerator here, not a decode accelerator.

GPU clock note (measured, no guesswork): the Mali-G610 is an integrated GPU sharing the board's DDR. The devfreq governor (simple_ondemand, stock) boosts it to 1 GHz under load and parks it at 200 MHz idle — confirmed by sampling cur_freq at 400 ms during the run (233/263 samples at 1000 MHz, temps 41→56 °C, no thermal trip). Leave the governor alone: forcingperformance/min/max via sysfs was counterproductive (idle readbacks that look like the clock collapsed, and it destabilized perfectly good runs). The numbers above are at the stock governor's real 1 GHz boost.

Results: Gemma 4 E4B (4B) — text GPU via the web flavor, vision GPU straight from stock

E4B (the 4B sibling) ships as three files in the HF repo: the stock gemma-4-E4B-it.litertlm (text+vision+audio, 3.66 GB) plus an -gpu and a -web (text-only, 2.97 GB) variant. On this 8 GB board the stock file cannot run on the GPU — two structural walls, both measured:

panthor job watchdog: a fixed one-shot pipeline dispatch (dmesg: job timeout ... seqno=144, same job on every E4B attempt) exceeds the driver's compiled-in 1 s job timeout even at the real 1 GHz boost (E2B's equivalent is seqno=77 and fits). The GPU is reset, Dawn reports VK_ERROR_DEVICE_LOST, generation aborts.

8 GB RAM ceiling: GPU buffers (pinned, unrescalable shmem) peak near 4–5 GB on top of the 3.66 GB model; every run ended in a global OOM-kill (EXIT=137) during iteration 2 even with a 2048-token KV cap + ringbuffers + disk cache.

The -web flavor dodges both — its finer op layout splits the killer dispatch under the 1 s bar and shrinks the GPU working set (peak 4.3 GB observed, holds 1 GHz the whole run):

Gemma 4 E4B lane Flavor Prefill (tk/s) Decode (tk/s) TTFT (s)
CPU (XNNPACK, 8 threads) stock 52.6 5.58 19.6
GPU (patched panvk) web (text-only) 75.1 4.54 13.9

(reproducible to the decimal across --runs 2; init 6.2 s; note --cache disk does not persist for the web flavor — expect a recompile each run. --speculative-decoding true gave no gain:4.37 vs 4.54 tok/s decode (this build ships no draft model).) Reference: Raspberry Pi 5 (16 GB, CPU) 51/3.2/20.5 — this board matches or beats it on every column. Same shape as E2B: GPU wins prefill +1.4× and TTFT 19.6→13.9 s, but decode stays CPU-favored (4.5 vs 5.6 tk/s).

E4B vision (VL) on GPU works — cpu text + gpu vision

The stock E4B's vision encoder is a separate 1477-op subgraph that fits under the panthor watchdog and inside 8 GB even though the full stock model's text lane does not. Drive it via the Engine API (this repo's vl_bench.py --model ...), LLM lane = CPU, vision encoder = GPU (patched panvk): the test image (resized to 912×672 → 2394 patches, near E4B's max_num_patches 2520) is encoded on the Mali.

E4B VL (LLM lane = cpu) vision cpu vision gpu (patched panvk)
TTFT (s) 12.3–12.5 9.6
Decode (tok/s) 7.33–7.37 7.28–7.39
Reply (same image) "coyote walking on a dirt path" identical

Text-only CPU control: TTFT ≈ 2.0 s (7-token prompt, decode ~7.8 tok/s). Isolated vision-encode cost: ~10.4 s CPU vs ~7.6 s GPU — the Mali cuts E4B vision latency by ~2.9 s/image (~27%). No watchdog hits, no OOM (peaks well below ceiling; vision weights are a fraction of a GB). Same conclusion as E2B: for vision work keep the LLM on CPU and let the GPU run the encoder.

# 1. SET ENVIRONMENT FOR GPU VISION DELEGATE
  export LITERT_VISION_DELEGATE=gpu

  # 2. RUN VL HARNESS FOR STOCK GEMMA-4-E4B
~/.local/share/uv/tools/litert-lm/bin/python vl_bench.py \
    --text cpu \
    --vision gpu \
    --image \
    --iters 3 \
    --model gemma-4-E4B-it.litertlm
# Ensure directory exists and download model
mkdir -p ~/litert-lm-bench
curl -L -o ~/litert-lm-bench/gemma-4-E4B-it-web.litertlm \
  https://huggingface.co/litert-community/gemma-4-E4B-it-litert-lm/resolve/main/gemma-4-E4B-it-web.litertlm

# Execute benchmark with standard LiteRT-LM flags
litert-lm benchmark ~/litert-lm-bench/gemma-4-E4B-it-web.litertlm \
  --backend=gpu \
  -p 1024 \
  -d 256 \
  --runs 2

Results: VL (vision) path

The benchmark sub-command only benches the text path. To time the vision encoder, drive the same engine via the Python API (vl_bench.py in this directory) — it loads the model, streams a real image (test_multi.jpg, httpbin.org/image/jpeg) + text through vision_backend=cpu|gpu and prints TTFT (vision encode + prefill + first token), per-chunk decode rate, and the reply. One sentence prompt, 13-token answer, ~290-token total prefill (280 vision + text), generations deterministic (temperature=0). Medians over 3+ runs after compile:

Combo (text / vision) VL TTFT (s) Decode (tok/s) Reply matches CPU?
cpu / cpu 8.93–9.40 17.4 baseline
cpu / gpu (patched panvk) 6.16–6.62 17.2 yes, identical
gpu / gpu (patched panvk) 5.62–6.16 9.8 yes, identical

Text-only controls (7-token prompt): CPU TTFT ≈ 0.8 s, GPU TTFT ≈ 0.57 s.

Isolating the vision cost (VL TTFT − text-only TTFT for the same text backend): ~8.1 s on a CPU vision encoder vs ~5.4 s GPU — the Mali GPU cuts vision-encoder latency by ~2.7–3 s per image (~30%), and the full-GPU combo is ~37% faster end-to-end (5.6 vs 8.9 s). The patched panvk covers the model's separate VISION_ENCODER subgraph too (no extra Dawn limits tripped).

Caveat: for short generations (< ~20 tokens) GPU decode is slower than CPU (9.8 vs 17.4 tok/s) — per-iteration GPU sync overhead — so for chatty/VL replies the CPU text lane with GPU vision is often the best mix (6.2 s TTFT, CPU-fast decode).

Benchmarking the VL path yourself

# 1. FETCH SAMPLE TEST IMAGE
  curl -sL -o ~/litert-lm-bench/test_multi.jpg https://httpbin.org/image/jpeg

  # 2. SET PYTHON INTERPRETER PATH
  VL_PY=~/.local/share/uv/tools/litert-lm/bin/python

  # 3. BASELINE: TEXT (CPU) + VISION (CPU)
  $VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --vision cpu --image --iters 3

  # 4. HYBRID: TEXT (CPU) + VISION (GPU) — EXPORT PANVK MESA ENVS FIRST
  export BUILD=~/litert-lm-bench/mesa-26.1.2/build-panvk
  export VK_ICD_FILENAMES="$BUILD/src/panfrost/vulkan/panfrost_devenv_icd.aarch64.json"
  export LD_LIBRARY_PATH="$BUILD/src/panfrost/vulkan:$HOME/litert-lm-bench/prebuilt_v0171"

  $VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --vision gpu --image --iters 3

  # 5. FULL ACCELERATION: TEXT (GPU) + VISION (GPU)
  $VL_PY ~/litert-lm-bench/vl_bench.py --text gpu --vision gpu --image --iters 3

  # 6. TEXT-ONLY CONTROLS (ISOLATE VISION-ENCODER COST)
  $VL_PY ~/litert-lm-bench/vl_bench.py --text cpu --iters 3
  $VL_PY ~/litert-lm-bench/vl_bench.py --text gpu --iters 3

Prints per-iteration TTFT / decode tok/s / reply. Run cases sequentially — parallel GPU runs contend for the Mali GPU and inflate the numbers (observed 2× TTFT when two ran at once).

6. Inference

# 1. DIRECT CLI EVALUATION 
  litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
    --cache disk --prompt "What is the capital of France?"

  # 2. INTERACTIVE REPL MODE
  litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm

  # 3. OPENAI-COMPATIBLE API SERVER
  litert-lm serve ~/litert-lm-bench/gemma-4-E2B-it.litertlm

Vision (and audio) input uses run with attachments — one per --attachment, placed before the first user text (images and audio can be mixed):

litert-lm run ~/litert-lm-bench/gemma-4-E2B-it.litertlm \
  --attachment ~/litert-lm-bench/test_multi.jpg \
  --prompt "What is in this image? Answer in one sentence." \
  --vision-backend cpu          
  # or gpu (patched panvk, §5) — ~2.7–3 s faster vision encode

--vision-backend/--audio-backend pick the encoder lane independently of --backend (which chooses the LLM lane). Like --backend gpu, --vision-backend gpu needs the §5 env (VK_ICD_FILENAMES + LD_LIBRARY_PATH) exported for that run.

Notes:

--cache disk persists compiled artifacts next to the model — the first GPU load compiles shaders, later loads start instantly. The load time is NOT paid on every run. Observed caches: <model>_*_mldrift_program_cache.bin (compiled kernels) and<model>_*_mldrift_weight_cache.bin (GPU-rearranged weights, ~0.8 GB). The vision encoder's kernels live in the same cache files — first GPU --vision-backend gpu run compiles, later
runs reuse.

GPU (patched panvk) is ~2.8× faster at prefill and ~2.7× on TTFT, but ~30% slower at decode — pick the lane by workload (§7.4).

Vision: --vision-backend gpu cuts vision-encode latency by ~2.7–3 s per image; CPU text + GPU vision is the best blend for short/chatty replies (§5 VL results).

Speculative decoding (--speculative-decoding true) can lift decode on CPU+GPU for rewrite/summarize/coding style prompts (Gemma 4 E2B supports it).

7. GPU deep dive: the blocker and how it was unblocked here

Everything below was reproduced with LiteRT-LM v0.17.1 (CLI + litert-lm-api), Mesa 26.1.2 (trixie-backports), kernel 7.2.4-edge-rockchip64 (mainline panthor).

7.1 The panvk limit blocker — RESOLVED with a one-line local patch

LiteRT's WebGPU accelerator uses Dawn, which enforces the WebGPU spec minimums against the Vulkan driver and refuses to proceed when they are unmet. Stock panvk reports:

maxImageDimension3D = 512      # WebGPU spec requires ≥ 2048

Dawn logs exactly this (PhysicalDeviceVk.cpp:794: "Insufficient Vulkan limits for maxTextureDimension3D ... must be at least 2048"), and stock panvk then dies with a null-pointer dispatch (SIGSEGV, pc=0x0) inside the WebGPU path on the first real GPU execution. The root cause is the driver limit, not packaging: the crash is identical even with the WebGPU prebuilts (libLiteRtTopKWebGpuSampler.so + friends) fetched from the litert-lm repo prebuilt/linux_arm64 @ v0.17.1 on LD_LIBRARY_PATH.

The fix (validated on this board): panvk caps 3D textures at 512 for Valhall
(PAN_ARCH <= 10), but the Mali-G610 hardware handles 2048³ — raising the advertised limit to 2048 makes Dawn accept the adapter:

-.maxImageDimension3D = PAN_ARCH <= 10 ? (1 << 9) : (1 << 14),
+.maxImageDimension3D = (1 << 11),       /* 2048 — meets WebGPU min (was 512 on arch <= 10) */

Why it's safe: Dawn only validates the advertised limit against its spec minimum, and the value also caps future image allocations — so if a kernel ever genuinely requested a 2048³ 3D texture, panvk would fail cleanly at allocation instead of corrupting anything. LiteRT/ML-Drift's LLM and vision-encoder kernels allocate buffers and 2D textures only, so real GPU behaviour is unchanged; Dawn simply stops rejecting the adapter, the SIGSEGV disappears, and the full benchmark runs (§5). Confirmed via vulkaninfo on the patched build: maxImageDimension3D = 2048. All other WebGPU minimums (per-stage descriptors, workgroup sizes, buffer ranges) were already comfortably met.

The build is process-scoped (VK_ICD_FILENAMES + LD_LIBRARY_PATH per run), the system Mesa is never touched, and reverting is unsetting two variables. Unless/until Mesa raises this limit upstream for Valhall, keep this build around for LiteRT-LM GPU runs.

7.2 The vendor route (Rockchip libmali) — needs a different kernel

Rockchip's proprietary blob would satisfy Dawn (it exposes Vulkan 1.3 with full limits), but the blob's userspace talks to Rockchip's proprietary kbase kernel driver (/dev/mali) and cannot attach to the mainline panthor driver:

"libmali-valhall-g610-g13p0-gbm"      → loads, exports NO Vulkan ICD (0 vk_* symbols)
"libmali-valhall-g610-g24p0-gbm"
  libMaliVulkan.so.1 (api 1.3.276)    → loads, "No mali devices found" then SIGSEGV (no /dev/mali)

Installable packages exist (tsukumijima/libmali-rockchip releases, e.g.
v1.9-1-20260312-bd33ee2), but they all presume the vendor kernel.

Unblock: flash a vendor-kernel image (kernel with kbase/mali.ko, e.g. Armbian images with the Rockchip BSP kernel), then install
libmali-valhall-g610-g13p0-gbm (or g24p0) and point Dawn/Vulkan at it
(VK_DRIVER_FILES/LD_LIBRARY_PATH). That same blob also provides OpenCL 3.0 (libMaliOpenCL.so + /etc/OpenCL/vendors/mali.icd), which would additionally satisfy the "would OpenCL be faster?" curiosity — for running LLMs, both WebGPU/Vulkan and OpenCL on the same GPU land in the same class of throughput; the win is the mature compiler's fast startup, not raw speed.

7.3 OpenCL on the current kernel — not available

Mesa's rusticl on panthor currently exposes zero devices (clinfo -l shows only the rusticl platform, no device). So there is no OpenCL device at all on the mainline stack, and LiteRT-LM has no Linux OpenCL accelerator anyway.

7.4 Practical advice for this board today

  • Chat / streaming (decode-bound): CPU is the better lane — 12.3 vs 8.7 tk/s decode.
  • Long prompts / RAG / document Q&A (prefill-bound): GPU pays off — 345 vs 124 tk/s prefill,
    3.1 vs 8.4 s TTFT.

Images / VL: --vision-backend gpu cuts vision-encode latency by ~2.7–3 s (~30%). Shortest VL TTFT is text+vision both on GPU (5.6 vs 8.9 s all-CPU), but keep the LLM on CPU when replies are short and chatty (6.2 s TTFT plus CPU-fast decode).

  • --speculative-decoding true can lift decode further on both lanes (Gemma 4 E2B supports it).

Serving an OpenAI-compatible endpoint (litert-lm serve)

litert-lm serve exposes an OpenAI API on 0.0.0.0:9379, so LAN clients (LiteCode, opencode, aider, Continue, scripts) use the board like any OpenAI endpoint. Import a model once (litert-lm import ./gemma-4-E2B-it.litertlm, registry at ~/.litert-lm/models/), then litert-lm serve. Per-model settings come from ~/.litert-lm/config.json (default + models.<id>), so clients stay plain-OpenAI:

{
  "default": { "backend": "cpu", "cpu_thread_count": 8, "cache": "disk" },
  "models": { "gemma-4-E2B-it.litertlm": { "speculative_decoding": true, "max_num_tokens": 4096 } }
}

Endpoints: GET /v1/models and POST /v1/chat/completions (streaming + non-streaming). OpenAI tools / tool_choice are bridged to the model's function-calling template — verified to return valid tool_calls JSON (non-streaming and streamed) and to complete the full agent cycle (assistant tool_call → tool result → final answer). /v1/embeddings exists but is not tied to LLM models.

Even though litert-lm describe reports Supports Function Call: NO for Gemma 4, the serve layer still bridges tools into the model template (empirically verified); that flag only means the CLI's interactive run has no native FC template.

RAM budget (E2B, CPU): max_num_tokens preallocates KV/ringbuffer arenas up front.Measured process RSS: 4096 → ~3.1 GB (4.9 GB free — recommended), 8192 → ~7 GB (1.1 GB free — works but leaves no headroom). Keep 4096 and let clients budget. Config keys (0.15+): backend, vision_backend, audio_backend, cpu_thread_count, cache,
max_num_tokens, speculative_decoding, thinking, thinking_budget,
gpu_decode_steps_per_sync, sampling. CLI flags override config. For a persistent border, run under systemd.

Verified client on this board

LiteCode (razvanneculai/litecode) — works end-to-end. Its Planner
({"synthesis","tasks"} strict JSON) and Executor (raw file content, no fences) prompts fit the 4096 budget; single-request task runs completed correctly and files linter-clean.
litecode.json:

{
  "provider": { "baseURL": "http://<yourIP>:9379/v1", "apiKey": "", "model": "gemma-4-E2B-it.litertlm" },
  "tokenLimit": 4096, "reservedOutputTokens": 1500, "systemPromptBudget": 1000, "maxParallelExecutors": 1
}

Troubleshooting

FATAL ERROR: This binary was compiled with dotprod enabled... → CPU lacks FEAT_DotProd, not supported by the ARM64 prebuilt wheels. RK3588 has it.

GPU benchmark errors but CPU works → with stock Mesa this is the §7.1 Dawn limit crash (maxImageDimension3D = 512); with the patched build, check vulkaninfo --summary again — panvk must list the Mali device and report maxImageDimension3D = 2048. ExportVK_ICD_FILENAMES+LD_LIBRARY_PATH for every run that should use the patched driver.

Could not load shared library libLiteRtTopKWebGpuSampler.so → the CLI wheel does not bundle the WebGPU samplers; fetch them from the litert-lm repo prebuilt/linux_arm64/ at the matching version tag (e.g. v0.17.1) and export their directory on LD_LIBRARY_PATH.

First GPU run is slow (one-time kernel compilation) — normal; use --cache disk so later loads skip it. Caches persist next to the model.

Bigger models (e.g. stock gemma-4-E4B-it.litertlm, 3.66 GB) fail on the GPU even with the patched driver: panthor ... job timeout in dmesg (fixed 1 s driver watchdog, not tunable; the model has a one-shot dispatch that exceeds it even at the real 1 GHz boost) theVK_ERROR_DEVICE_LOST, and/or a global OOM-kill (EXIT=137) when pinned GPU buffers + the model outgrow 8 GB. E2B fits, E4B-stock doesn't. Use the text-only -web flavor (gemma-4-E4B-it-web.litertlm, 2.97 GB) for E4B-on-GPU — its finer kernels stay under the watchdog and its GPU set peaks ~4.3 GB (§5 E4B section).

Join the conversation on the Armbian Forum and let us know how it runs on your hardware!

View the full article

Account

Navigation

Search

Search

Configure browser push notifications

Chrome (Android)
  1. Tap the lock icon next to the address bar.
  2. Tap Permissions → Notifications.
  3. Adjust your preference.
Chrome (Desktop)
  1. Click the padlock icon in the address bar.
  2. Select Site settings.
  3. Find Notifications and adjust your preference.