Describe the bug
A delete task crashes VictoriaLogs within milliseconds when logs are being ingested at the same time:
panic VictoriaLogs/lib/logstorage/datadb.go:1065 BUG: unexpected number of parts removed; got 1, want 10
Reproduced on v1.52.0 and on current master (3119079) with the test below: it fails every run, in well under a second.
Cause
datadb.deleteRows (v1.52.0 datadb.go:1499) claims parts from a snapshot without checking that they are still live:
getPartsForTimeRange snapshots the parts (incRef only).
hasMatchingRows scans each part without holding partsLock.
- Under
partsLock, it claims any part with isInMerge == false.
In between, a regular merge can take the same part, finish, and remove it from ddb.inmemoryParts / smallParts / bigParts in swapSrcWithDstParts. After that, mustMergePartsInternal's deferred releasePartsToMerge sets isInMerge = false on the retired part. So in step 3 the retired part looks free. deleteRows claims it and merges it. Then swapSrcWithDstParts removes fewer parts than it expects, and hits the logger.Panicf.
Delete tasks run one at a time (watchDeleteTasks), so concurrent deletes are not needed. Ingestion alone keeps in-memory parts flushing and merging, which opens the window.
To Reproduce
Add lib/logstorage/datadb_delete_race_test.go:
package logstorage
import (
"sync"
"testing"
"time"
)
// TestDatadbDeleteRowsConcurrentWithMerges runs deleteRows while other goroutines keep
// ingesting and flushing. deleteRows claims parts it snapshotted earlier; if a concurrent
// merge retired one of them in between, swapSrcWithDstParts panics with
// "BUG: unexpected number of parts removed".
func TestDatadbDeleteRowsConcurrentWithMerges(t *testing.T) {
path := t.Name()
s := newTestStorage()
mustCreatePartition(path)
pt := mustOpenPartition(s, path)
lrSample := newTestLogRows(3, 5, 0)
var tenantIDs []TenantID
for _, sid := range lrSample.streamIDs {
tenantIDs = append(tenantIDs, sid.tenantID)
}
PutLogRows(lrSample)
sso := newTestStorageSearchOptions(tenantIDs, &filterNoop{}, nil)
deadline := time.Now().Add(5 * time.Second)
var wg sync.WaitGroup
for w := range 4 {
wg.Add(1)
go func() {
defer wg.Done()
for i := 0; time.Now().Before(deadline); i++ {
lr := newTestLogRows(3, 5, 0)
pt.mustAddRows(lr)
PutLogRows(lr)
if (i+w)%4 == 0 {
pt.debugFlush()
}
}
}()
}
stopCh := make(chan struct{})
for time.Now().Before(deadline) {
pt.deleteRows(sso, stopCh)
}
wg.Wait()
mustClosePartition(pt)
mustDeletePartition(path)
closeTestStorage(s)
}
go test ./lib/logstorage/ -run TestDatadbDeleteRowsConcurrentWithMerges -count=1
v1.52.0, 5 of 5 runs:
panic: BUG: unexpected number of parts removed; got 2, want 7 [recovered, repanicked]
panic: BUG: unexpected number of parts removed; got 6, want 8 [recovered, repanicked]
panic: BUG: unexpected number of parts removed; got 2, want 4 [recovered, repanicked]
panic: BUG: unexpected number of parts removed; got 6, want 13 [recovered, repanicked]
panic: BUG: unexpected number of parts removed; got 1, want 4 [recovered, repanicked]
master 3119079 fails the same way (datadb.go:1069). The stack is the same as in production: deleteRows → mustMergePartsInternal → swapSrcWithDstParts.
Suggested fix
Claim only parts that are still registered; otherwise ask for a repeat, as deleteRows already does for parts that are busy in a merge. With this change on master the test passes 5 of 5 runs, and go test ./lib/logstorage/ passes:
@@ -1514,7 +1515,10 @@ func (ddb *datadb) deleteRows(pso *partitionSearchOptions, stopCh <-chan struct{
}
ddb.partsLock.Lock()
- if !pw.isInMerge {
+ if !ddb.containsPartLocked(pw) {
+ // A concurrent merge has already replaced pw with a new part; the new part must be processed on the next run.
+ needRepeat = true
+ } else if !pw.isInMerge {
pw.isInMerge = true
pwsToMerge = append(pwsToMerge, pw)
} else {
@@ -1540,3 +1544,7 @@ func appendAllPartsForMergeLocked(dst, src []*partWrapper) []*partWrapper {
}
return dst
}
+
+func (ddb *datadb) containsPartLocked(pw *partWrapper) bool {
+ return slices.Contains(ddb.inmemoryParts, pw) || slices.Contains(ddb.smallParts, pw) || slices.Contains(ddb.bigParts, pw)
+}
(plus "slices" in the imports). The linear scan is per matching part under partsLock. A flag set on retired parts in swapSrcWithDstParts would avoid the scan.
Version
$ docker run --rm victoriametrics/victoria-logs:v1.52.0 --version
victoria-logs-20260716-022232-tags-v1.52.0-0-g46a54c976f
Logs
=== RUN TestDatadbDeleteRowsConcurrentWithMerges
2026-09-23T15:13:36.432Z info VictoriaMetrics/lib/memory/memory.go:45 limiting caches to 41231686041 bytes, leaving 27487790695 bytes to the OS according to
-memory.allowedPercent=60, system memory limit 68719476736 bytes
2026-09-23T15:13:36.437Z panic /tmp/vl-repro/lib/logstorage/datadb.go:1065 BUG: unexpected number of parts removed; got 2, want 6
--- FAIL: TestDatadbDeleteRowsConcurrentWithMerges (0.01s)
panic: BUG: unexpected number of parts removed; got 2, want 6 [recovered, repanicked]
goroutine 40 [running]:
testing.tRunner.func1.2({0x101426e40, 0x9d01622090})
/Users/matthewadams/go/pkg/mod/golang.org/toolchain@v0.0.1-go1.26.5.darwin-arm64/src/testing/testing.go:1974 +0x1a0
testing.tRunner.func1()
/Users/matthewadams/go/pkg/mod/golang.org/toolchain@v0.0.1-go1.26.5.darwin-arm64/src/testing/testing.go:1977 +0x318
panic({0x101426e40?, 0x9d01622090?})
/Users/matthewadams/go/pkg/mod/golang.org/toolchain@v0.0.1-go1.26.5.darwin-arm64/src/runtime/panic.go:860 +0x12c
github.com/VictoriaMetrics/VictoriaMetrics/lib/logger.logMessageInternal({0x100f01eba, 0x5}, {0x9d017160c0, 0x36}, {0x9d012c01b0, 0x2b})
/tmp/vl-repro/vendor/github.com/VictoriaMetrics/VictoriaMetrics/lib/logger/logger.go:329 +0x6fc
github.com/VictoriaMetrics/VictoriaMetrics/lib/logger.logLevelSkipframes(0x9d01482640?, {0x100f01eba, 0x5}, {0x100f46b94, 0x38}, {0x9d014099f8, 0x2, 0x2})
/tmp/vl-repro/vendor/github.com/VictoriaMetrics/VictoriaMetrics/lib/logger/logger.go:162 +0x2d8
github.com/VictoriaMetrics/VictoriaMetrics/lib/logger.logLevel(...)
/tmp/vl-repro/vendor/github.com/VictoriaMetrics/VictoriaMetrics/lib/logger/logger.go:147
github.com/VictoriaMetrics/VictoriaMetrics/lib/logger.Panicf(...)
/tmp/vl-repro/vendor/github.com/VictoriaMetrics/VictoriaMetrics/lib/logger/logger.go:143
github.com/VictoriaMetrics/VictoriaLogs/lib/logstorage.(*datadb).swapSrcWithDstParts(0x9d01482640, {0x9d01394140, 0x6, 0x9d015be1e0?}, 0x0, 0x0)
/tmp/vl-repro/lib/logstorage/datadb.go:615 +0x738
github.com/VictoriaMetrics/VictoriaLogs/lib/logstorage.(*datadb).deleteRows(0x9d01482640, 0x9d012bc060, 0x9d0139ae70)
/tmp/vl-repro/lib/logstorage/datadb.go:1524 +0x2b0
github.com/VictoriaMetrics/VictoriaLogs/lib/logstorage.(*partition).deleteRows(0x9d01395340, 0x9d013973b0, 0x9d0139ae70)
/tmp/vl-repro/lib/logstorage/partition.go:281 +0x50
github.com/VictoriaMetrics/VictoriaLogs/lib/logstorage.TestDatadbDeleteRowsConcurrentWithMerges(0x9d014ae488?)
/tmp/vl-repro/lib/logstorage/datadb_delete_race_test.go:45 +0x224
testing.tRunner(0x9d014ae488, 0x1014bdcf8)
/Users/matthewadams/go/pkg/mod/golang.org/toolchain@v0.0.1-go1.26.5.darwin-arm64/src/testing/testing.go:2036 +0xc4
created by testing.(*T).Run in goroutine 1
/Users/matthewadams/go/pkg/mod/golang.org/toolchain@v0.0.1-go1.26.5.darwin-arm64/src/testing/testing.go:2101 +0x3a8
FAIL github.com/VictoriaMetrics/VictoriaLogs/lib/logstorage 0.326s
FAIL
Screenshots
No response
Used command-line flags
Container defaults.
Additional information
#825 / #845 were a different delete-API panic (block_stream_merger.go minTimestamp), fixed in v1.39.0.
Describe the bug
A delete task crashes VictoriaLogs within milliseconds when logs are being ingested at the same time:
Reproduced on v1.52.0 and on current
master(3119079) with the test below: it fails every run, in well under a second.Cause
datadb.deleteRows(v1.52.0 datadb.go:1499) claims parts from a snapshot without checking that they are still live:getPartsForTimeRangesnapshots the parts (incRef only).hasMatchingRowsscans each part without holdingpartsLock.partsLock, it claims any part withisInMerge == false.In between, a regular merge can take the same part, finish, and remove it from
ddb.inmemoryParts/smallParts/bigPartsinswapSrcWithDstParts. After that,mustMergePartsInternal's deferredreleasePartsToMergesetsisInMerge = falseon the retired part. So in step 3 the retired part looks free.deleteRowsclaims it and merges it. ThenswapSrcWithDstPartsremoves fewer parts than it expects, and hits thelogger.Panicf.Delete tasks run one at a time (
watchDeleteTasks), so concurrent deletes are not needed. Ingestion alone keeps in-memory parts flushing and merging, which opens the window.To Reproduce
Add
lib/logstorage/datadb_delete_race_test.go:v1.52.0, 5 of 5 runs:
master3119079 fails the same way (datadb.go:1069). The stack is the same as in production:deleteRows→mustMergePartsInternal→swapSrcWithDstParts.Suggested fix
Claim only parts that are still registered; otherwise ask for a repeat, as
deleteRowsalready does for parts that are busy in a merge. With this change onmasterthe test passes 5 of 5 runs, andgo test ./lib/logstorage/passes:(plus
"slices"in the imports). The linear scan is per matching part underpartsLock. A flag set on retired parts inswapSrcWithDstPartswould avoid the scan.Version
Logs
Screenshots
No response
Used command-line flags
Container defaults.
Additional information
#825 / #845 were a different delete-API panic (block_stream_merger.go minTimestamp), fixed in v1.39.0.