This ticket has been redacted with ChatGPT
Description
After replacing a failed disk in a RAID1 array, ZimaOS kept reporting that the RAID still needed recovery even though mdadm reported the array as fully healthy.
The issue was caused by stale/inconsistent data in:
/var/lib/casaos/db/local-storage.db
The raids table still reported:
shortage = 1
disk_number = 3
devices = /dev/sdc /dev/sdd /dev/sde
while the actual RAID had only 2 members and was healthy.
There was also an incorrect/stale mapping of physical disks to /dev/sdX device names in:
/var/lib/casaos/db/local-storage.json
This caused ZimaOS to sometimes display a disk belonging to another RAID as if it belonged to this array.
Environment
ZimaOS Local Storage build:
- git commit:
06979747e87d8c8ff7ffcae850e9a332eb241563
- build date:
2026-08-07T05:36:04Z
RAIDs:
md0: RAID1, 2x NVMe
md1: RAID1, 2x EMTEC SSD
md2: RAID1, 2x Crucial MX500 4TB
Filesystem on top of mdadm: Btrfs
Actual mdadm state
After fixing the RAID geometry, mdadm reported:
md2 : active raid1 sdc[0] sdd[2]
3906886464 blocks super 1.2 [2/2] [UU]
And:
Raid Level : raid1
Raid Devices : 2
Total Devices : 2
State : clean
Active Devices : 2
Working Devices : 2
Failed Devices : 0
Spare Devices : 0
So the array itself was fully healthy.
ZimaOS database state
However, this command:
sqlite3 -header -column /var/lib/casaos/db/local-storage.db 'SELECT * FROM raids;'
showed for Main-Storage:
name shortage devices disk_number status
Main-Storage 1 /dev/sdc /dev/sdd /dev/sde 3 ok
The members field only contained 2 actual Crucial MX500 drives, so the database was internally inconsistent:
disk_number = 3
- only 2 actual members
Incorrect device mapping
local-storage.json also contained stale /dev/sdX assignments.
For example, the physical drives had changed names after reboot/device replacement, but ZimaOS still associated their serial numbers with previous /dev/sdX paths.
This resulted in situations where the UI could show a disk from another RAID as belonging to md2.
This is especially problematic because /dev/sdX names are not persistent identifiers and may change after reboot or hardware changes.
Manual workaround
I stopped the Local Storage service:
systemctl stop zimaos-local-storage.service
Verified all mdadm arrays were still healthy:
Then manually corrected only the Main-Storage row:
sqlite3 /var/lib/casaos/db/local-storage.db "UPDATE raids SET shortage=0, devices='/dev/sdc /dev/sdd', disk_number=2 WHERE id=3;"
After restarting the service:
systemctl start zimaos-local-storage.service
the ZimaOS UI immediately showed the RAID correctly.
After reboot, the filesystem also mounted normally again instead of read-only.
Expected behavior
After a RAID repair/replacement:
-
ZimaOS should re-read the actual mdadm metadata.
-
disk_number should reflect the real number of RAID devices.
-
shortage should become 0 when mdadm reports [2/2] [UU].
-
Device membership should be tracked using persistent identifiers such as:
- serial number
- mdadm UUID
- WWN
/dev/disk/by-id
rather than relying on /dev/sdX names.
Additional observation
At service startup, the logs also showed:
and:
failed to load bays from global setting
Get "http://127.0.0.1:38710/v2/settings/global/fe.custom.setting.storage.bay":
dial tcp 127.0.0.1:38710: connect: connection refused
It is not clear whether these are related, but they may be worth checking.
Impact
This is more than a cosmetic UI issue.
Because ZimaOS believed the array still had a missing member, it continued to offer RAID recovery even though mdadm considered the RAID healthy.
The stale /dev/sdX mapping is also potentially dangerous because disks from different RAID arrays may be displayed incorrectly after device enumeration changes.
This ticket has been redacted with ChatGPT
Description
After replacing a failed disk in a RAID1 array, ZimaOS kept reporting that the RAID still needed recovery even though
mdadmreported the array as fully healthy.The issue was caused by stale/inconsistent data in:
/var/lib/casaos/db/local-storage.dbThe
raidstable still reported:shortage = 1disk_number = 3devices = /dev/sdc /dev/sdd /dev/sdewhile the actual RAID had only 2 members and was healthy.
There was also an incorrect/stale mapping of physical disks to
/dev/sdXdevice names in:/var/lib/casaos/db/local-storage.jsonThis caused ZimaOS to sometimes display a disk belonging to another RAID as if it belonged to this array.
Environment
ZimaOS Local Storage build:
06979747e87d8c8ff7ffcae850e9a332eb2415632026-08-07T05:36:04ZRAIDs:
md0: RAID1, 2x NVMemd1: RAID1, 2x EMTEC SSDmd2: RAID1, 2x Crucial MX500 4TBFilesystem on top of mdadm: Btrfs
Actual mdadm state
After fixing the RAID geometry,
mdadmreported:And:
So the array itself was fully healthy.
ZimaOS database state
However, this command:
sqlite3 -header -column /var/lib/casaos/db/local-storage.db 'SELECT * FROM raids;'showed for
Main-Storage:The
membersfield only contained 2 actual Crucial MX500 drives, so the database was internally inconsistent:disk_number = 3Incorrect device mapping
local-storage.jsonalso contained stale/dev/sdXassignments.For example, the physical drives had changed names after reboot/device replacement, but ZimaOS still associated their serial numbers with previous
/dev/sdXpaths.This resulted in situations where the UI could show a disk from another RAID as belonging to
md2.This is especially problematic because
/dev/sdXnames are not persistent identifiers and may change after reboot or hardware changes.Manual workaround
I stopped the Local Storage service:
Verified all mdadm arrays were still healthy:
Then manually corrected only the
Main-Storagerow:sqlite3 /var/lib/casaos/db/local-storage.db "UPDATE raids SET shortage=0, devices='/dev/sdc /dev/sdd', disk_number=2 WHERE id=3;"After restarting the service:
the ZimaOS UI immediately showed the RAID correctly.
After reboot, the filesystem also mounted normally again instead of read-only.
Expected behavior
After a RAID repair/replacement:
ZimaOS should re-read the actual mdadm metadata.
disk_numbershould reflect the real number of RAID devices.shortageshould become0when mdadm reports[2/2] [UU].Device membership should be tracked using persistent identifiers such as:
/dev/disk/by-idrather than relying on
/dev/sdXnames.Additional observation
At service startup, the logs also showed:
and:
It is not clear whether these are related, but they may be worth checking.
Impact
This is more than a cosmetic UI issue.
Because ZimaOS believed the array still had a missing member, it continued to offer RAID recovery even though
mdadmconsidered the RAID healthy.The stale
/dev/sdXmapping is also potentially dangerous because disks from different RAID arrays may be displayed incorrectly after device enumeration changes.