Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
237 changes: 237 additions & 0 deletions pkg/backingstore/backingstore.go
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@ func Cmd() *cobra.Command {
CmdList(),
CmdReconcile(),
CmdRunRemovePendingPods(),
CmdReplace(),
)
return cmd
}
Expand Down Expand Up @@ -352,6 +353,27 @@ func CmdDelete() *cobra.Command {
return cmd
}

// CmdReplace returns a CLI command
func CmdReplace() *cobra.Command {
cmd := &cobra.Command{

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this replace work on all the buckets at once? So if there are 100 buckets and we are doing migration on those buckets will it create any issue?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, it updates all affected tiers in a single system_store.make_changes() call — one atomic bulk DB write. Even with 100 buckets, it's just updating tier documents, not moving data. The heavy lifting (actual data replication in migration mode) is handled by the existing mirror_writer background worker which processes buckets independently. So the replace call itself is fast regardless of bucket count.

Use: "replace <old-backing-store> <new-backing-store>",
Short: "Replace one backing store with another across all buckets",
Long: `Replace references to one backing store with another across all bucket tiers and account defaults.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think, we should also provide buckets specific replace too,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, may be we can add --bucket-name flag (mandatory)

@kajalpareek-lab kajalpareek-lab Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought about this, but I don’t think it works cleanly because tiers can be shared by multiple buckets through tiering policies.
For example, if we replace a pool in a tier for Bucket A, Bucket B could also be affected if both buckets use the same tier. So adding per-bucket filtering could be misleading because users may think they are changing only one bucket when they could actually impact others.
Also, the main reason this bug exists (DFBUGS-6233) is that there was no way to update all non-OBC buckets at once. Making --bucket-name mandatory would mean users have to manually provide every bucket, which is exactly the problem we are trying to solve.
If there is a real customer need to filter by bucket, we can add it as an optional flag later if there's a concrete use case. But we should design it carefully because of the shared-tier behavior.


Use --migrate to first enable mirroring between the old and new backing stores,
allowing the system to replicate existing data before completing the switch.

Workflow:
1. noobaa backingstore replace <old> <new> --migrate (start mirroring)
2. Wait for data replication to complete
3. noobaa backingstore replace <old> <new> (finalize replacement)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why we need this step?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The --migrate step is optional — it's only needed when buckets have existing data on the old backing store that you don't want to lose. It adds the new pool as a mirror so the mirror_writer background service copies existing objects to the new pool before you cut over.

If your buckets are empty or you don't care about existing data, skip --migrate entirely and just run noobaa backingstore replace directly. That does the swap in one step.

The two-step flow exists for production scenarios where losing existing data is not acceptable.

4. oc delete backingstore <old> (remove the old store)`,
Run: RunReplace,
}
Comment thread
aayushchouhan09 marked this conversation as resolved.
cmd.Flags().Bool("migrate", false, "Enable mirroring to replicate data before replacement")
return cmd
}

// CmdStatus returns a CLI command
func CmdStatus() *cobra.Command {
cmd := &cobra.Command{
Expand Down Expand Up @@ -945,6 +967,221 @@ func RunCreatePVPool(cmd *cobra.Command, args []string) {
}

// RunDelete runs a CLI command
// RunReplace runs a CLI command to replace one backing store with another
func RunReplace(cmd *cobra.Command, args []string) {
log := util.Logger()

if len(args) != 2 || args[0] == "" || args[1] == "" {
log.Fatalf(`❌ Missing expected arguments: <old-backing-store> <new-backing-store> %s`, cmd.UsageString())
}

oldBSName := args[0]
newBSName := args[1]

if oldBSName == newBSName {
log.Fatalf(`❌ Old and new backing store names must be different`)
}

migrate, _ := cmd.Flags().GetBool("migrate")

// Verify both backing stores exist in Kubernetes
oldBS := util.KubeObject(bundle.File_deploy_crds_noobaa_io_v1alpha1_backingstore_cr_yaml).(*nbv1.BackingStore)
oldBS.Name = oldBSName
oldBS.Namespace = options.Namespace
if !util.KubeCheck(oldBS) {
log.Fatalf(`❌ BackingStore %q not found in namespace %q`, oldBSName, options.Namespace)
}

newBS := util.KubeObject(bundle.File_deploy_crds_noobaa_io_v1alpha1_backingstore_cr_yaml).(*nbv1.BackingStore)
newBS.Name = newBSName
newBS.Namespace = options.Namespace
if !util.KubeCheck(newBS) {
log.Fatalf(`❌ BackingStore %q not found in namespace %q`, newBSName, options.Namespace)
}

if newBS.Status.Phase != nbv1.BackingStorePhaseReady {
log.Fatalf(`❌ BackingStore %q is not Ready (current phase: %s)`, newBSName, newBS.Status.Phase)
}

nbClient := system.GetNBClient()

// Verify both pools exist in NooBaa and new pool is healthy
_, err := nbClient.ReadPoolAPI(nb.ReadPoolParams{Name: oldBSName})
if err != nil {
log.Fatalf(`❌ Pool %q not found in NooBaa: %s`, oldBSName, err)
}
newPoolInfo, err := nbClient.ReadPoolAPI(nb.ReadPoolParams{Name: newBSName})
if err != nil {
log.Fatalf(`❌ Pool %q not found in NooBaa: %s`, newBSName, err)
}
if newPoolInfo.Mode != "OPTIMAL" {
log.Fatalf(`❌ Pool %q is not healthy (mode: %s)`, newBSName, newPoolInfo.Mode)
}

if migrate {
log.Infof("🔄 Starting migration: mirroring data from %q to %q", oldBSName, newBSName)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are we really starting the migration here? Is it done by background worker?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We're not starting any new process. The RPC just adds the new pool as a separate mirror group in each tier and sets data_placement: MIRROR. The existing mirror_writer background service (in src/server/bg_services/mirror_writer.js) already runs continuously — it detects MIRROR tiers and copies chunks between mirror groups automatically. So we're just configuring tiers in a way the existing worker picks up.

log.Infof(" This adds %q as a mirror to all tiers currently using %q", newBSName, oldBSName)
log.Infof(" The background mirror_writer will replicate existing data automatically")
} else {
log.Infof("🔄 Replacing %q with %q across all bucket tiers and account defaults", oldBSName, newBSName)
}

systemInfo, err := nbClient.ReadSystemAPI()
if err != nil {
log.Fatalf(`❌ Failed to read system info: %s`, err)
}

// Build tier name → TierInfo lookup from system info
tierByName := map[string]nb.TierInfo{}
for _, t := range systemInfo.Tiers {
tierByName[t.Name] = t
}

// Phase 1: Update all tiers that reference the old pool
tiersUpdated := replacePoolInTiers(nbClient, systemInfo.Buckets, tierByName, oldBSName, newBSName, migrate)

// Phase 2: Update account default_resource references
accountsUpdated := replacePoolInAccounts(nbClient, systemInfo.Accounts, oldBSName, newBSName)

if tiersUpdated == 0 && accountsUpdated == 0 {
log.Infof("ℹ️ No tiers or accounts reference %q — nothing to replace", oldBSName)
return
}

if migrate {
log.Infof("✅ MIRROR_STARTED")
} else {
log.Infof("✅ REPLACED")
}
log.Infof(" Tiers updated: %d", tiersUpdated)
log.Infof(" Accounts updated: %d", accountsUpdated)

log.Infof("")
if migrate {
log.Infof("📋 Next steps:")
log.Infof(" 1. Wait for data replication to complete")
log.Infof(" Check progress: noobaa bucket status <bucket-name>")
Comment on lines +1062 to +1063

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we have any tracker or something to verify the mirroring status?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not a dedicated one today. The user can check noobaa bucket status to see the mirror configuration is in place, and the mirror_writer logs progress to the NooBaa server logs.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you please share an example of how bucket status would show mirror status (screenshot may be)?

And now we can move to step 2.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After --migrate, noobaa bucket status will show the tier with both pools listed under MIRROR placement. The tier mode reflects the mirror configuration. Once mirror_writer finishes copying all chunks, both pools show OPTIMAL status. At that point you run the finalize step.

No single '% complete' indicator today — the user checks that both pools are OPTIMAL in bucket status. We can add a progress indicator as a follow-up.

log.Infof(" 2. Finalize the replacement:")
log.Infof(" noobaa backingstore replace %s %s", oldBSName, newBSName)
Comment on lines +1062 to +1065

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you explain this how --mirror works? Do we need to run the replace command twice here?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, it's two steps:

noobaa backingstore replace old new --migrate adds new as a mirror alongside old. The background mirror_writer then copies existing data from old to new automatically.
noobaa backingstore replace old new (without --migrate) removes old and keeps only new.
The two-step approach makes sure data is fully replicated before we cut over. If you don't care about existing data (e.g., empty buckets), you can skip step 1 and just run the direct replace.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why we are not doing it in one step?

Are we manually checking whether the mirroring is done or not. Will step 2 check this part and remove old one after mirroring is done.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two steps because the user needs to decide when replication is done we can't automate that judgment. Step 1 sets up mirroring then the mirror_writer starts copying data in the background. The user checks replication progress and decides when to finalize. Step 2 (finalize) just swaps the pools it doesn't check replication status, it trusts the user has verified it. Making it one step would risk data loss if we cut over before all data is replicated. The two-step approach is the safe path.

} else {
log.Infof("📋 The old backing store %q has been detached from all tiers.", oldBSName)
}
Comment thread
aayushchouhan09 marked this conversation as resolved.
log.Infof(" To clean up the old backing store:")
log.Infof(" oc patch noobaa/noobaa -n %s --type json --patch='[{\"op\":\"add\",\"path\":\"/spec/manualDefaultBackingStore\",\"value\":true}]'", options.Namespace)
log.Infof(" oc delete backingstore %s -n %s", oldBSName, options.Namespace)
}

// replacePoolInTiers iterates all buckets from system info, looks up each tier
// in the pre-built map, and replaces references to oldPool with newPool.
// In migrate mode it adds newPool as a mirror group alongside oldPool.
func replacePoolInTiers(nbClient nb.Client, buckets []nb.BucketInfo, tierByName map[string]nb.TierInfo, oldPool, newPool string, migrate bool) int {
log := util.Logger()
tiersUpdated := 0
seen := map[string]bool{}

for _, bucket := range buckets {
if bucket.Tiering == nil {
continue
}
for _, tierItem := range bucket.Tiering.Tiers {
tierName := tierItem.Tier
if seen[tierName] {
continue
}
seen[tierName] = true

tierInfo, ok := tierByName[tierName]
if !ok {
log.Warnf("⚠️ Tier %q not found in system info, skipping", tierName)
continue
}

if !util.Contains(tierInfo.AttachedPools, oldPool) {
continue
}

updateParams := buildTierUpdate(tierInfo, oldPool, newPool, migrate)
if updateParams == nil {
continue
}
err := nbClient.UpdateTierAPI(*updateParams)
if err != nil {
log.Fatalf(`❌ Failed to update tier %q: %s`, tierName, err)
}
log.Infof(" Updated tier %q (bucket %q)", tierName, bucket.Name)
tiersUpdated++
}
}
return tiersUpdated
}

// buildTierUpdate constructs the UpdateTierParams to swap oldPool with newPool.
// In migrate mode, it adds newPool alongside the existing pools (MIRROR).
// In replace mode, it replaces oldPool with newPool and preserves data_placement.
func buildTierUpdate(tier nb.TierInfo, oldPool, newPool string, migrate bool) *nb.UpdateTierParams {
if migrate {
// Add newPool to attached_pools for mirroring; keep oldPool
if util.Contains(tier.AttachedPools, newPool) {
return nil // already mirrored
}
pools := make([]string, 0, len(tier.AttachedPools)+1)
pools = append(pools, tier.AttachedPools...)
pools = append(pools, newPool)
return &nb.UpdateTierParams{
Name: tier.Name,
DataPlacement: "MIRROR",
AttachedPools: pools,
}
}

// Replace mode: swap oldPool → newPool, preserve data_placement
pools := make([]string, 0, len(tier.AttachedPools))
for _, p := range tier.AttachedPools {
if p == oldPool {
if !util.Contains(pools, newPool) {
pools = append(pools, newPool)
}
} else {
pools = append(pools, p)
}
}

placement := tier.DataPlacement
// If only one pool remains, MIRROR is meaningless — use SPREAD
if len(pools) <= 1 && placement == "MIRROR" {
placement = "SPREAD"
}
return &nb.UpdateTierParams{
Name: tier.Name,
DataPlacement: placement,
AttachedPools: pools,
}
}

// replacePoolInAccounts updates any account whose default_resource points to oldPool.
func replacePoolInAccounts(nbClient nb.Client, accounts []nb.AccountInfo, oldPool, newPool string) int {
log := util.Logger()
updated := 0
for i := range accounts {
acct := &accounts[i]
if acct.DefaultResource != oldPool {
continue
}
newRes := newPool
err := nbClient.UpdateAccountS3Access(nb.UpdateAccountS3AccessParams{
Email: acct.Email,
S3Access: acct.HasS3Access,
DefaultResource: &newRes,
})
if err != nil {
log.Fatalf(`❌ Failed to update account %q default_resource: %s`, acct.Email, err)
}
log.Infof(" Updated account %q default_resource", acct.Email)
updated++
}
return updated
}

func RunDelete(cmd *cobra.Command, args []string) {

log := util.Logger()
Expand Down
7 changes: 7 additions & 0 deletions pkg/nb/api.go
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ type Client interface {
CreateCloudPoolAPI(CreateCloudPoolParams) error
UpdateCloudPoolAPI(UpdateCloudPoolParams) error
CreateTierAPI(CreateTierParams) error
UpdateTierAPI(UpdateTierParams) error
CreateNamespaceResourceAPI(CreateNamespaceResourceParams) error
CreateTieringPolicyAPI(TieringPolicyInfo) error

Expand Down Expand Up @@ -295,6 +296,12 @@ func (c *RPCClient) CreateTieringPolicyAPI(params TieringPolicyInfo) error {
return c.Call(req, nil)
}

// UpdateTierAPI calls tier_api.update_tier()
func (c *RPCClient) UpdateTierAPI(params UpdateTierParams) error {
req := &RPCMessage{API: "tier_api", Method: "update_tier", Params: params}
return c.Call(req, nil)
}

// DeleteBucketAPI calls bucket_api.delete_bucket()
func (c *RPCClient) DeleteBucketAPI(params DeleteBucketParams) error {
req := &RPCMessage{API: "bucket_api", Method: "delete_bucket", Params: params}
Expand Down
49 changes: 49 additions & 0 deletions pkg/nb/api_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -46,3 +46,52 @@ func TestBigIntUnmarshalBig(t *testing.T) {
func TestBigIntUnmarshalBigNoPeta(t *testing.T) {
testBigIntUnmarshal(t, `{"n":99}`)
}

func TestUpdateTierParamsMarshal(t *testing.T) {
params := UpdateTierParams{
Name: "tier1",
DataPlacement: "SPREAD",
AttachedPools: []string{"pool-a", "pool-b"},
}
data, err := json.Marshal(params)
if err != nil {
t.Fatal(err)
}
expected := `{"name":"tier1","data_placement":"SPREAD","attached_pools":["pool-a","pool-b"]}`
if string(data) != expected {
t.Fatalf("unexpected marshal result: %s", string(data))
}
}

func TestUpdateTierParamsMarshalPoolsOnly(t *testing.T) {
params := UpdateTierParams{
Name: "tier1",
AttachedPools: []string{"new-pool"},
}
data, err := json.Marshal(params)
if err != nil {
t.Fatal(err)
}
expected := `{"name":"tier1","attached_pools":["new-pool"]}`
if string(data) != expected {
t.Fatalf("unexpected marshal result: %s", string(data))
}
}

func TestTierInfoUnmarshal(t *testing.T) {
tierJSON := `{"name":"tier1","data_placement":"SPREAD","attached_pools":["pool-a","pool-b"]}`
tier := TierInfo{}
err := json.Unmarshal([]byte(tierJSON), &tier)
if err != nil {
t.Fatal(err)
}
if tier.Name != "tier1" {
t.Fatalf("expected name=tier1, got %s", tier.Name)
}
if tier.DataPlacement != "SPREAD" {
t.Fatalf("expected data_placement=SPREAD, got %s", tier.DataPlacement)
}
if len(tier.AttachedPools) != 2 {
t.Fatalf("expected 2 attached_pools, got %d", len(tier.AttachedPools))
}
}
7 changes: 7 additions & 0 deletions pkg/nb/types.go
Original file line number Diff line number Diff line change
Expand Up @@ -566,6 +566,13 @@ type TierItem struct {
Mode string `json:"mode,omitempty"`
}

// UpdateTierParams is the params of tier_api.update_tier()
type UpdateTierParams struct {
Name string `json:"name"`
DataPlacement string `json:"data_placement,omitempty"`
AttachedPools []string `json:"attached_pools,omitempty"`
}

// DeleteBucketParams is the params of bucket_api.delete_bucket()
type DeleteBucketParams struct {
Name string `json:"name"`
Expand Down
1 change: 1 addition & 0 deletions test/cli/test_cli_flow.sh
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ function post_install_tests {
test_noobaa_cr_deletion
test_noobaa_loadbalancer_source_subnet
test_multinamespace_bucketclass
test_backingstore_replace
check_default_backingstore #It deletes all the buckets and non-default accounts, creates new backingstore and attach it to default admin account and then deletes default backingstore
}

Expand Down
Loading
Loading