September 16Sep 16 I have an interesting discovery that smells suspiciously like a kernel bug. I find it easiest to see this bug by running btop which shows a historical per core cpu usage graph so it's easy to see when one stops being used. You can see the same thing in ordinary top but it's less easy to see than with btop. I have a Radxa Rock 5B with an RK3588 chip that I have been stress testing to make sure the hardware is OK before I start to use it in anger. As part of that I ran `openssl speed -multi 8` to spin up 8 threads using 100% cpu each. On all other machines that I've ever done that with it runs all X cores at 100% until it runs out of tests or is Ctrl-C'ed. Watching top on the Rock5B shows something very odd happens. When it first starts it correctly seems to use all 8 cores at 100% but after a while, maybe 20s, maybe a couple of minutes, it suddenly just stops using one core, sometimes more than one core. This behaviour is also true if running `stress --cpu 8` which I used to make sure this was not an openssl bug, same symptoms, 8x100% cpu for a while then entire cores go completely idle. I have also run these tests on an Odroid HC1 which uses a similar big:little set of cores (Samsung Exynos 5422) as the Rock 5B - 4 smaller slower low power cores, 4 higher power faster cores. That test was run on plain Debian 13 with a 6.12.107 kernel which is why I ran a test on the Rock 5B using armbian 6.12.58 kernel to see if this was a kernel regression. On the Odroid it just works and needs no special actions to make it use all 8 cores simultaneously. I left it to run for the full 1h+ that the openssl speed run takes to complete. It also works correctly on an rpi5 though I only ran that with 4 threads to match its cores. If I change my test to pin a separate process to individual cores using `for core in {0..7}; do taskset -c $core openssl speed & done` then it quite happily runs all 8 of them on all cores until they finish so this appears to me to show that it's not a hardware problem as all 8 processes run on all 8 cores for the entire hour or so it takes for an openssl speed test to complete. I have tested this with various Armbian kernels starting with 6.18.44 then 7.1.8 and 7.2.5 from the beta channel and also reverting to 6.12.58. All show the same symptoms - it just abandons running tasks on one physical core. I have also tested with higher numbers of threads (9-12) and the problem still exhibits itself but takes longer, the more threads, the longer it takes to happen but it does happen eventually. If I run the same 8 process version of the test but omit the taskset -c $core so the system can schedule the task where it likes then it also shows the idle core problem. I am suspecting a kernel scheduler bug since when I manually pin a process to a core so that each one is effectively dedicated to running just that one task then the problem does not occur. I have asked a couple of other Rock 5B users and they also have the same symptoms. I've also seen the symptoms on mine extend to 2 idle cores. These cores are inactive even when all openssl/stress processes show that they are ready to run in the output of ps fax. I do not think this is temperature related as my Rock 5B was consistently reporting itself at less than 57C during these tests. This is what sar reports for openssl/stress with 8 threads: 21:00:51 CPU %user %nice %system %iowait %steal %idle 21:10:43 all 87.39 0.00 0.11 0.00 0.00 12.50 21:20:51 all 87.39 0.00 0.11 0.00 0.00 12.50 21:30:51 all 87.39 0.00 0.11 0.00 0.00 12.50 21:40:43 all 87.38 0.00 0.11 0.00 0.00 12.50 21:50:51 all 87.39 0.00 0.11 0.00 0.00 12.50 and for the taskset -c $core version of the test: 16:26:20 CPU %user %nice %system %iowait %steal %idle 16:30:20 all 7.72 0.00 0.21 11.54 0.00 80.52 16:40:20 all 99.63 0.00 0.37 0.00 0.00 0.00 16:50:20 all 99.81 0.00 0.19 0.00 0.00 0.00 17:00:22 all 99.82 0.00 0.18 0.00 0.00 0.00 17:10:22 all 99.85 0.00 0.15 0.00 0.00 0.00 17:20:20 all 99.86 0.00 0.14 0.00 0.00 0.00 17:30:20 all 99.52 0.00 0.14 0.04 0.00 0.30 17:40:22 all 99.60 0.00 0.40 0.00 0.00 0.00 17:50:20 all 99.83 0.00 0.17 0.00 0.00 0.00 18:00:20 all 99.83 0.00 0.17 0.00 0.00 0.00 The first sample there will be where it was not running for the entire 10 minute period. I've posted similar on the linux-rockchip mailing list and have received a reply from someone at rock-chips.com asking me to rebuild the kernel with CONFIG_SCHED_CLUSTER=n so I'd really like to try to find the deb-src package for linux-image-edge-rockchip64 so that I can rebuild it with that option (not) set. Edit: this appears to be fixed in the 7.3.0 kernel series, specifically tested -rc3. Edited September 18Sep 18 by TrevorH Found the fix
September 17Sep 17 Providing logs with armbianmonitor -u helps with troubleshooting and significantly raises chances that issue gets addressed.
September 17Sep 17 Can't confirm your observation on Radxa ROCK 5 ITX, but my kernel is build with # CONFIG_SCHED_CLUSTER is not set
September 17Sep 17 I could reproduce on my NanoPi-R6C (RK3588S). It currently does nothing special other than showing an internal home automation website in a room where no-one really is. So I thought make it even more dedicated by doing first: 'systemctl isolate multi-user.target' so no firefox or so that can disturb. Then 2 ssh terminal sessions, 1 with 'btop' and 1 with 'openssl speed -multi 8' Ran for more than an hour, tempurature went up from 31 to 50 Celsius but still constant 8x 100% load. kernel: 7.3.0-rc2-bleedingedge-rockchip64 #1 SMP PREEMPT Sun Sep 6 22:07:20 UTC 2026 aarch64 GNU/Linux U-Boot: U-Boot SPL 2026.01_armbian-2026.01-S127a-Paca4-He05b-V36ee-B5da4-R448a (Aug 14 2026 - 13:26:35 +0000) I discovered I had not configured sysstat properly, not on this test computer, but also not on another server. I fixed it there, it was basically setting it to "true" in /etc/default/sysstat Then I did scp that file to the test computer still running the openssl test. BUT, when switching to its btop terminal, the 2 lowest numbered cores were idle and was kept like that. Until I restarted the openssl test, then again all 8 cores 100%. So my perception is that that some considerable action like a remote file copy disturbs the scheduling and it does not recover. My NanoPi-R6C has its Btrfs formatted rootfs on NVME. Another remarkable thing was that the date/time of the test computer was not correct, I see chrony did not correct it, it likely has to do with the brute-force switch from graphical to multi-user. Restart chronyd, then OK again; But don't think this impact scheduling issue discussed in this topic. /boot# grep CONFIG_SCHED_CLUSTER config-7.3.0-rc2-bleedingedge-rockchip64 CONFIG_SCHED_CLUSTER=y I could maybe run it also on my ROCK5B, uses other kernel and bootloader, but cannot do too much testing as it is my main server running 24/7 also running some KVM instances that probably make it too specific non-reproducible as I use CPU pinning there. Edited September 17Sep 17 by eselarm
September 17Sep 17 After some thinking, I thought checking various kernel configs of various installs/distros: grep CONFIG_SCHED_CLUSTER /boot/config-* config-6.1.115-vendor-rk35xx:# CONFIG_SCHED_CLUSTER is not set config-6.1.172-vendor-rk35xx:# CONFIG_SCHED_CLUSTER is not set config-6.12.107+deb13-arm64:# CONFIG_SCHED_CLUSTER is not set config-6.18.44-current-rockchip64:CONFIG_SCHED_CLUSTER=y config-6.18.50+rpt-rpi-v8:CONFIG_SCHED_CLUSTER=y config-7.1.8-edge-rockchip64:CONFIG_SCHED_CLUSTER=y config-7.1.13+deb14-arm64:CONFIG_SCHED_CLUSTER=y config-7.1.13+deb14-arm64-16k:CONFIG_SCHED_CLUSTER=y config-7.2.4-1-default:CONFIG_SCHED_CLUSTER=y config-7.3.0-rc2-bleedingedge-rockchip64:CONFIG_SCHED_CLUSTER=y Maybe I use Armbian Build to just build latest 7.3 kernel with CONFIG_SCHED_CLUSTER=n, I need to read docs first as changing the .config did appear a bit confusing to me w.r.t. steps to take, but it should work. On the other hand, I have no clue what the effect is on let's say power-consumption, on average and also for various use-cases. For the NanoPi-R6C I am trying to get lowest idle power, which is still almost double for mainline kernels compared to vendor 6.1 last time I measured. Highest load on my NanoPi-R6C is building kernels and/or images, is anyway a constant high rate of starting and stopping processes (12 gcc threads/tasks by default) and the semi-permanent idling case for 1 or 2 cores would not happen.
September 17Sep 17 Maybe an interesting thing is to do: # for i in 0 1 2 3 4 5 6 7 ; do cat /sys/devices/system/cpu/cpu$i/topology/cluster_id ; done 0 0 0 0 1 1 2 2 This is on RK3588 (and where the kernels have CONFIG_SCHED_CLUSTER=y All other 4-core SoC's I see the same for all 4 core, can be '-1' or '36' depending on bootloader and kernel. For RK3588 it reminds me of something I saw elsewhere in the DeviceTree or so AFAIR, is 3 power or clocking domains, but check yourself. Also Radxa Q6A would be interesting for listing this as well as the AllWinner 2xA76+6xA55 and other bigLITTLE SoC's.
September 17Sep 17 Author I am suspecting this is a bug in the linux kernel scheduler that only gets triggered on machines with an asymmetrical cpu configuration. I'm pretty sure that it will always happen on the Rock 5B as I've had reports from at least 2 other 5B users to say they can reproduce this. It would be interesting to see any results from people running machines with a mix of big:little cores that are not the rk3588 to see if the problem is more widespread (I suspect it is).
September 17Sep 17 2 hours ago, TrevorH said: I am suspecting this is a bug in the linux kernel scheduler that only gets triggered on machines with an asymmetrical cpu configuration. Sofar I have not seen any testrun on a RK3588 with CONFIG_SCHED_CLUSTER=n, I don not know what is compiled/configured when it is not set. The other ROCK5B users likely have not used a kernel with CONFIG_SCHED_CLUSTER=n. I am not 100% sure that this is a bug; you want to see all cores loaded, but it does not mean that the total will be faster. It also depends on cache structures and levels and how clusters are interacting and are connected. You can have a look at Ampere SoC's versus SoC's from Intel. It might be that just not involving slow cores will result in better/higher cache hits for faster cores, just a thought. It might be that if you want a very long run with just dedicated CPU-bound tasks, that you actually should not do scheduling. So use other specific kernel, such that you anyhow get better results. It is more like a DSP then or sort of HW accelerator. I have been designing with- and using Cortex-R variants and also traditional DSP's, maybe that is why I think this.
September 17Sep 17 Author Well I've just rebuilt the 6.18.52 current kernel without CONFIG_SCHED_CLUSTER set and have rebooted into it and am now testing with stress --cpu 8. So far it's using all 8 cores but it's not been running for long enough to prove it yet. Will update later once it's run and broken/not/broken. root@rock-5b:~# grep CONFIG_SCHED_CL /boot/config-6.18.52-current-rockchip64 # CONFIG_SCHED_CLASS_EXT is not set # CONFIG_SCHED_CLUSTER is not set Edit: 15 minutes in and it's still using all 8 cores at 100% so this does seem to be a bug in the CONFIG_SCHED_CLUSTER code path. It's never lasted for as long as this without pinning processes to specific cores. Edit2: 40 minutes into stress --cpu 8 and it's still using all 8 cores at 100% so I think that proves there is a bug in the SCHED_CLUSTER stuff. Since stress never ends, I'm going to kill that now as 40 minutes sounds like enough. Edited September 17Sep 17 by TrevorH
September 18Sep 18 Author So I ran `openssl speed -multi 8 rsa md5 sha256 sha512 aes sha1 camellia des rmd160` on 2 builds of the 6.18.52 kernel, one from the repos and one self-built with CONFIG_SCHED_CLUSTER=n and then ran diff against the final 21 lines of output where the stats are. You can see that the all of the md5 tests and 4 of the 6 sha1 tests are almost the same. These are run at the start of the run and I notice that the distro kernel runs all 8 cores for the first 20-30s so that explains why those first results are more or less the same. After that it starts using only 7 cores and the numbers diverge more widely. I checked a few and mostly there seems to be an 8-11% advantage in favour of the CONFIG_SCHED_CLUSTER=n version (i.e non-distro kernel). So CLUSTER=n results are the lefthand side and CLUSTER=y are on the right. trevor@rock-5b:~$ diff -y -W 190 openssl-speed-nocluster.txt openssl-speed-cluster.txt md5 206646.70k 682082.69k 1627733.76k 2515694.25k 3003411.11k 3047533.23k | md5 207893.48k 681828.80k 1624703.66k 2514480.47k 3003370.15k 3047145.47k sha1 257265.18k 945564.91k 2826501.12k 5732223.32k 8417888.94k 8727172.44k | sha1 257001.97k 944134.21k 2818819.67k 5729986.56k 8258613.03k 7898929.54k rmd160 168775.57k 482578.01k 1047489.96k 1490576.73k 1703146.84k 1721171.97k | rmd160 157569.82k 446793.28k 963480.75k 1360659.80k 1549350.23k 1565089.82k sha256 256593.18k 949375.79k 2850635.69k 5781916.67k 8441995.26k 8740929.54k | sha256 241875.40k 894078.82k 2667497.15k 5343714.37k 7705458.01k 7961350.38k sha512 139462.71k 558478.93k 1124027.14k 1811831.47k 2209688.23k 2246688.77k | sha512 130139.93k 520255.22k 1040071.77k 1665936.38k 2023230.12k 2055602.18k des-cbc 0.00 0.00 0.00 0.00 0.00 0.00 des-cbc 0.00 0.00 0.00 0.00 0.00 0.00 des-ede3 132712.38k 138286.40k 139840.77k 140261.44k 140372.65k 140245.64k | des-ede3 122223.67k 126929.64k 128361.86k 128859.14k 128809.72k 128849.24k aes-128-cbc 2887098.42k 6649536.64k 10139807.06k 11986084.86k 12734630.57k 12803970.39k | aes-128-cbc 2737165.32k 6205807.32k 9220290.30k 10736674.05k 11360862.21k 11410365.80k aes-192-cbc 2751396.09k 5874383.27k 8473588.05k 9655262.55k 10161837.40k 10203824.13k | aes-192-cbc 2609197.06k 5462457.41k 7717889.27k 8693462.87k 9127941.46k 9163740.50k aes-256-cbc 2674905.48k 5340891.67k 7373585.83k 8255187.97k 8579080.19k 8609371.48k | aes-256-cbc 2531961.38k 4973017.60k 6722701.40k 7453166.55k 7728204.46k 7748107.85k camellia-128-cbc 582227.33k 679838.68k 710770.01k 721328.47k 723978.92k 724489.56 | camellia-128-cbc 537508.58k 621512.36k 648128.60k 655762.20k 658434.73k 658625.88 camellia-192-cbc 468416.18k 530513.96k 550038.44k 556270.25k 557826.05k 558000.81 | camellia-192-cbc 433255.91k 484146.24k 500183.47k 506022.57k 507199.49k 507164.29 camellia-256-cbc 467775.22k 530535.47k 550066.77k 556213.59k 556876.29k 557940.74 | camellia-256-cbc 430679.27k 483797.80k 500581.66k 505632.77k 507202.22k 507150.34 sign verify encrypt decrypt sign/s verify/s encr./s decr./s sign verify encrypt decrypt sign/s verify/s encr./s decr./s rsa 512 bits 0.000016s 0.000001s 0.000002s 0.000017s 64304.1 755378.7 645816.8 57525.3 | rsa 512 bits 0.000017s 0.000001s 0.000002s 0.000019s 58896.7 690687.0 594987.3 52892.2 rsa 1024 bits 0.000083s 0.000004s 0.000005s 0.000086s 11990.6 233562.6 218317.7 11679.9 | rsa 1024 bits 0.000092s 0.000005s 0.000005s 0.000094s 10854.4 211381.1 198357.9 10589.4 rsa 2048 bits 0.000571s 0.000016s 0.000016s 0.000574s 1752.4 64301.8 62614.2 1743.0 | rsa 2048 bits 0.000633s 0.000017s 0.000018s 0.000636s 1580.0 57918.6 56515.5 1573.2 rsa 3072 bits 0.001760s 0.000034s 0.000035s 0.001765s 568.1 29361.9 28935.6 566.5 | rsa 3072 bits 0.001957s 0.000038s 0.000038s 0.001959s 511.0 26431.1 26051.8 510.3 rsa 4096 bits 0.003995s 0.000060s 0.000060s 0.004000s 250.3 16738.8 16563.8 250.0 | rsa 4096 bits 0.004442s 0.000066s 0.000067s 0.004449s 225.1 15041.4 14899.8 224.8 rsa 7680 bits 0.031106s 0.000207s 0.000208s 0.031131s 32.1 4839.0 4814.8 32.1 | rsa 7680 bits 0.034616s 0.000230s 0.000231s 0.034496s 28.9 4351.1 4323.8 29.0 rsa 15360 bits 0.189898s 0.000820s 0.000821s 0.184244s 5.3 1219.4 1217.3 5.4 | rsa 15360 bits 0.211541s 0.000911s 0.000917s 0.211495s 4.7 1097.8 1090.9 4.7
September 18Sep 18 I did 2 runs with kernel 7.3.0-rc3-bleedingedge-rockchip64: 1st compiled with 'CONFIG_SCHED_CLUSTER is not set' 2nd compiled with 'CONFIG_SCHED_CLUSTER=y' So command as root: openssl speed -multi 8 rsa md5 sha256 sha512 aes sha1 camellia des rmd160 > $(timestamp).txt 2>&1 # diff -y -W 190 20260918_101405.txt 20260918_131402.txt md5 195473.11k 649713.49k 1552381.95k 2399496.19k 2867533.14k 2907575.64k | md5 193680.95k 643423.83k 1545036.12k 2381227.69k 2787740.33k 2907422.72k sha1 244179.63k 895310.46k 2693081.17k 5312768.68k 7934184.11k 8319795.20k | sha1 238851.57k 889777.58k 2673776.90k 5448719.02k 7900075.35k 8331236.69k rmd160 160717.62k 459518.78k 998510.25k 1421213.35k 1624244.22k 1641054.21k | rmd160 158611.37k 457548.65k 995554.05k 1421269.33k 1626974.89k 1643954.18k sha256 244299.85k 904019.26k 2719061.16k 5515426.13k 8056171.18k 8333934.59k | sha256 240987.41k 894283.18k 2696836.69k 5502382.08k 8063797.93k 8348718.42k sha512 132576.50k 530496.13k 1072615.51k 1729068.37k 2108967.59k 2143742.63k | sha512 131973.71k 526479.34k 1071986.69k 1730009.43k 2112348.16k 2146364.07k des-cbc 0.00 0.00 0.00 0.00 0.00 0.00 des-cbc 0.00 0.00 0.00 0.00 0.00 0.00 des-ede3 126802.83k 131960.75k 133454.93k 133826.22k 133958.31k 133918.53k | des-ede3 126945.35k 132138.60k 133681.92k 134063.45k 134160.38k 134142.38k aes-128-cbc 2749082.15k 6362386.43k 9682607.36k 11434917.55k 12150382.59k 12211273.73k | aes-128-cbc 2760766.55k 6371641.15k 9690578.60k 11451418.28k 12167091.54k 12230066.18k aes-192-cbc 2626037.58k 5604254.76k 8093569.28k 9210415.45k 9697626.79k 9734498.99k | aes-192-cbc 2632188.38k 5580423.04k 8100798.89k 9222450.86k 9712006.49k 9748419.93k aes-256-cbc 2552706.54k 5102452.86k 7039338.75k 7876314.79k 8187565.40k 8212135.94k | aes-256-cbc 2558937.55k 5113163.04k 7050647.38k 7887683.93k 8199165.27k 8224549.55k camellia-128-cbc 557388.59k 649808.17k 678503.17k 688174.08k 690517.33k 690918.74 | camellia-128-cbc 558809.32k 650699.29k 679674.71k 689085.10k 691882.67k 691595.95 camellia-192-cbc 449286.66k 506755.63k 524808.96k 530548.39k 532198.74k 532239.70 | camellia-192-cbc 450000.65k 507649.49k 525672.28k 531326.29k 532916.91k 532938.75 camellia-256-cbc 447559.94k 506777.09k 524796.33k 530373.97k 532146.86k 531984.13 | camellia-256-cbc 448415.03k 507589.10k 525694.81k 531273.73k 532854.10k 532781.23 sign verify encrypt decrypt sign/s verify/s encr./s decr./s sign verify encrypt decrypt sign/s verify/s encr./s decr./s rsa 512 bits 0.000016s 0.000001s 0.000002s 0.000018s 61323.6 704891.6 604001.5 54820.9 | rsa 512 bits 0.000016s 0.000001s 0.000002s 0.000018s 61356.9 720433.6 607364.1 54805.3 rsa 1024 bits 0.000088s 0.000004s 0.000005s 0.000090s 11420.5 222529.8 206941.7 11123.6 | rsa 1024 bits 0.000087s 0.000004s 0.000005s 0.000090s 11434.6 222705.9 206928.2 11133.9 rsa 2048 bits 0.000599s 0.000016s 0.000017s 0.000602s 1670.0 61262.5 59550.0 1661.5 | rsa 2048 bits 0.000598s 0.000016s 0.000017s 0.000601s 1671.7 61327.8 59606.2 1662.6 rsa 3072 bits 0.001849s 0.000036s 0.000036s 0.001852s 540.9 27990.4 27534.4 539.8 | rsa 3072 bits 0.001847s 0.000036s 0.000036s 0.001850s 541.4 28020.8 27561.6 540.5 rsa 4096 bits 0.004196s 0.000063s 0.000063s 0.004203s 238.3 15939.6 15763.5 237.9 | rsa 4096 bits 0.004192s 0.000063s 0.000063s 0.004197s 238.5 15957.3 15777.2 238.3 rsa 7680 bits 0.032663s 0.000217s 0.000218s 0.032670s 30.6 4607.5 4582.3 30.6 | rsa 7680 bits 0.032621s 0.000217s 0.000218s 0.032637s 30.7 4612.2 4587.5 30.6 rsa 15360 bits 0.199505s 0.000862s 0.000863s 0.193183s 5.0 1160.7 1158.6 5.2 | rsa 15360 bits 0.199280s 0.000860s 0.000862s 0.192447s 5.0 1162.8 1159.9 5.2 This time I had my NanoPi-R6C set to default to multi-user.target. All done via remote ssh. I had started a btop session as well but kept 800% all the time in both runs. I won't look back to previous kernels, so for me I don't see a clear bug, as at least even with this very artificial computing session (for me) I see no significant difference and I anyhow have no clue where or what to change in kernel code (the dev/torwalds branch).
September 18Sep 18 Author I had not tested 7.3 as it's still in RC status but I have now and it looks like this bug is fixed there. The same openssl command ran for 8m22s and kept all 8 cores at 100% with CONFIG_SCHED_CLUSTER=y. The numbers from the 2 runs now differ by fractions of a %, the highest I found was 0.25% better on the CLUSTER=n older kernel than on 7.3 with CLUSTER=y but 0.25% is measurement error and more importantly I can see that it uses all 8 cores all of the time rather than just orphaning one at random.
September 19Sep 19 TrevorH said: Kinda wild that the fix landed upstream while you were busy rebuilding kernels lol, fingers crossed it gets backported to 6.18 so nobody has to run an rc just to get the 8th core back.
Join the conversation
You can post now and register later. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.