Skip to content

Changes needed for updatign TF to version 2.21 - #51713

Merged
cmsbuild merged 5 commits into
cms-sw:masterfrom
smuzaffar:tf2.21-changes
Aug 31, 2026
Merged

Changes needed for updatign TF to version 2.21#51713
cmsbuild merged 5 commits into
cms-sw:masterfrom
smuzaffar:tf2.21-changes

Conversation

@smuzaffar

Copy link
Copy Markdown
Contributor

Tensorflow 2.21.0 related changes. These are back ported from CMSSW_20_1_TF_X branch used by special TF_X IBs

akritkbehera and others added 4 commits August 15, 2026 08:36
TF 2.21 marks tensorflow::Status as deprecated (ABSL_DEPRECATE_AND_INLINE),
it is now an inline alias for absl::Status. Update PhysicsTools/TensorFlow,
RecoTauTag/RecoTau, and RecoTracker/PixelTrackFitting to use absl::Status
directly.

Also comment out an unused Eigen matrix (cm2) in RiemannFit.h whose only
assignment was already disabled.
@cmsbuild

cmsbuild commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

cms-bot internal usage

@cmsbuild

Copy link
Copy Markdown
Contributor

-code-checks

Logs: https://cmssdt.cern.ch/SDT/code-checks/cms-sw-PR-51713/50618

ERROR: Build errors found during clang-tidy run.

src/DQMServices/Core/interface/ROOTFilePB.pb.h:14:10: error: 'google/protobuf/runtime_version.h' file not found [clang-diagnostic-error]
   14 | #include "google/protobuf/runtime_version.h"
      |          ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
src/DQMServices/Core/interface/ROOTFilePB.pb.h:159:10: warning: annotate this function with 'override' or (rarely) 'final' [modernize-use-override]
--
gmake: *** [config/SCRAM/GMake/Makefile.coderules:129: code-checks] Error 2
gmake: *** [There are compilation/build errors. Please see the detail log above.] Error 2

@smuzaffar

Copy link
Copy Markdown
Contributor Author

code checks should be run with tool-conf generated via cms-sw/cmsdist#10789 . Let wait for cms-sw/cmsdist#10789 to generate the tool-conf

@smuzaffar

Copy link
Copy Markdown
Contributor Author

code-checks with cms.week0.PR_cb4a0836/100.0-fcdbd34291f19a2674c959cf8ff3ed7b

@cmsbuild

Copy link
Copy Markdown
Contributor

+code-checks

Logs: https://cmssdt.cern.ch/SDT/code-checks/cms-sw-PR-51713/50620

@cmsbuild

Copy link
Copy Markdown
Contributor

A new Pull Request was created by @smuzaffar for master.

It involves the following packages:

  • DQMServices/Core (dqm)
  • PhysicsTools/TensorFlow (ml)
  • PhysicsTools/TensorFlowAOT (ml)
  • RecoTauTag/RecoTau (reconstruction)
  • RecoTracker/PixelTrackFitting (reconstruction)

@Moanwar, @cmsbuild, @ctarricone, @gabrielmscampos, @hjkwon260, @jfernan2, @mandrenguyen, @rseidita, @srimanob, @valsdav, @y19y19 can you please review it and eventually sign? Thanks.
@GiacomoSguazzoni, @VinInn, @VourMa, @azotz, @barvic, @dgulhan, @elusian, @felicepantaleo, @gpetruc, @makortel, @mbluj, @mmasciov, @mmusich, @mtosi, @riga, @rovere this is something you requested to watch as well.
@ftenchini, @mandrenguyen, @sextonkennedy you are the release manager for this.

cms-bot commands are listed here

@smuzaffar

Copy link
Copy Markdown
Contributor Author

enable gpu

@smuzaffar

Copy link
Copy Markdown
Contributor Author

please test with cms-sw/cmsdist#10789

@cmsbuild

Copy link
Copy Markdown
Contributor

-1

Failed Tests: UnitTests
Size: This PR adds an extra 16KB to repository
Summary: https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-f8baaa/55510/summary.html
COMMIT: e639964
CMSSW: CMSSW_20_1_X_2026-08-23-0000/el9_amd64_gcc14
Additional Tests: GPU,AMD_MI300X,AMD_W7900,NVIDIA_H100,NVIDIA_L40S,NVIDIA_T4
User test area: For local testing, you can use /cvmfs/cms-ci.cern.ch/week0/cms-sw/cmssw/51713/55510/install.sh to create a dev area with all the needed externals and cmssw changes.

The following merge commits were also included on top of IB + this PR after doing git cms-merge-topic:

You can see more details here:
https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-f8baaa/55510/git-recent-commits.json
https://cmssdt.cern.ch/SDT/jenkins-artifacts/pull-request-integration/PR-f8baaa/55510/git-merge-result

Failed Unit Tests

I found 15 errors in the following unit tests:

---> test SagittaBiasNtuplizer had ERRORS
---> test PVValidation had ERRORS
---> test PrimaryVertex had ERRORS
and more ...

Comparison Summary

Summary:

  • You potentially added 279 lines to the logs
  • Reco comparison results: 19196 differences found in the comparisons
  • DQMHistoTests: Total files compared: 47
  • DQMHistoTests: Total histograms compared: 3836516
  • DQMHistoTests: Total failures: 31057
  • DQMHistoTests: Total nulls: 14
  • DQMHistoTests: Total successes: 3805427
  • DQMHistoTests: Total skipped: 18
  • DQMHistoTests: Total Missing objects: 0
  • DQMHistoSizes: Histogram memory added: -0.0030000000000000027 KiB( 46 files compared)
  • DQMHistoSizes: changed ( 2024.0030001 ): -0.004 KiB JetMET/SUSYDQM
  • DQMHistoSizes: changed ( 2024.0040001 ): 0.016 KiB JetMET/SUSYDQM
  • DQMHistoSizes: changed ( 2025.0000002 ): 0.020 KiB JetMET/SUSYDQM
  • DQMHistoSizes: changed ( 2025.0010001 ): -0.035 KiB JetMET/SUSYDQM
  • Checked 203 log files, 173 edm output root files, 47 DQM output files
  • TriggerResults: no differences found

AMD_MI300X Comparison Summary

There are some workflows for which there are errors in the baseline:
37634.402 step 2
37634.404 step 2
The results for the comparisons for these workflows could be incomplete
This means most likely that the IB is having errors in the relvals.The error does NOT come from this pull request

Summary:

  • You potentially removed 106 lines from the logs
  • ROOTFileChecks: Some differences in event products or their sizes found
  • Reco comparison results: 165 differences found in the comparisons
  • DQMHistoTests: Total files compared: 6
  • DQMHistoTests: Total histograms compared: 158030
  • DQMHistoTests: Total failures: 18850
  • DQMHistoTests: Total nulls: 0
  • DQMHistoTests: Total successes: 139180
  • DQMHistoTests: Total skipped: 0
  • DQMHistoTests: Total Missing objects: 0
  • DQMHistoSizes: Histogram memory added: 0.0 KiB( 5 files compared)
  • Checked 22 log files, 18 edm output root files, 6 DQM output files
  • TriggerResults: no differences found

AMD_W7900 Comparison Summary

There are some workflows for which there are errors in the baseline:
37634.404 step 2
The results for the comparisons for these workflows could be incomplete
This means most likely that the IB is having errors in the relvals.The error does NOT come from this pull request

Summary:

  • You potentially removed 32 lines from the logs
  • ROOTFileChecks: Some differences in event products or their sizes found
  • Reco comparison results: 130 differences found in the comparisons
  • DQMHistoTests: Total files compared: 7
  • DQMHistoTests: Total histograms compared: 173739
  • DQMHistoTests: Total failures: 19364
  • DQMHistoTests: Total nulls: 0
  • DQMHistoTests: Total successes: 154375
  • DQMHistoTests: Total skipped: 0
  • DQMHistoTests: Total Missing objects: 0
  • DQMHistoSizes: Histogram memory added: 0.0 KiB( 6 files compared)
  • Checked 24 log files, 20 edm output root files, 7 DQM output files
  • TriggerResults: no differences found

NVIDIA_H100 Comparison Summary

Summary:

NVIDIA_L40S Comparison Summary

Summary:

NVIDIA_T4 Comparison Summary

Summary:

  • No significant changes to the logs found
  • Reco comparison results: 130 differences found in the comparisons
  • DQMHistoTests: Total files compared: 7
  • DQMHistoTests: Total histograms compared: 173739
  • DQMHistoTests: Total failures: 18548
  • DQMHistoTests: Total nulls: 0
  • DQMHistoTests: Total successes: 155191
  • DQMHistoTests: Total skipped: 0
  • DQMHistoTests: Total Missing objects: 0
  • DQMHistoSizes: Histogram memory added: 0.0 KiB( 6 files compared)
  • Checked 25 log files, 20 edm output root files, 7 DQM output files
  • TriggerResults: no differences found

Max Memory Comparisons exceeding threshold

@cms-sw/core-l2 , I found 21 workflow step(s) with memory usage exceeding the error threshold:

Expand to see workflows ...
  • Error: Workflow 2025.0000002_RunZeroBias2025B_10k step2 max memory diff 53.9 exceeds +/- 30.0 MiB
  • Error: Workflow 2025.0000002_RunZeroBias2025B_10k step3 max memory diff 45.8 exceeds +/- 30.0 MiB
  • Error: Workflow 2025.0010001_RunJetMET02025C_10k step3 max memory diff 46.2 exceeds +/- 30.0 MiB
  • Error: Workflow 2025.0010001_RunJetMET02025C_10k step2 max memory diff 55.2 exceeds +/- 30.0 MiB
  • Error: Workflow 18434.0_TTbar_14TeV+2026 step3 max memory diff 46.0 exceeds +/- 30.0 MiB
  • Error: Workflow 18434.0_TTbar_14TeV+2026 step2 max memory diff 54.1 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.0_TTbar_14TeV+Run4D127 step5 max memory diff 44.6 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.0_TTbar_14TeV+Run4D127 step2 max memory diff 45.8 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.0_TTbar_14TeV+Run4D127 step3 max memory diff 46.1 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.75_TTbar_14TeV+Run4D127_HLT75e33Timing step2 max memory diff 50.4 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.911_TTbar_14TeV+Run4D127_DD4hep step2 max memory diff 45.8 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.911_TTbar_14TeV+Run4D127_DD4hep step3 max memory diff 46.1 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.911_TTbar_14TeV+Run4D127_DD4hep step5 max memory diff 44.6 exceeds +/- 30.0 MiB
  • Error: Workflow 37696.0_CloseByPGun_CE_E_Front_120um+Run4D127 step2 max memory diff 45.0 exceeds +/- 30.0 MiB
  • Error: Workflow 37696.0_CloseByPGun_CE_E_Front_120um+Run4D127 step5 max memory diff 44.5 exceeds +/- 30.0 MiB
  • Error: Workflow 37696.0_CloseByPGun_CE_E_Front_120um+Run4D127 step3 max memory diff 45.8 exceeds +/- 30.0 MiB
  • Error: Workflow 37700.0_CloseByPGun_CE_H_Coarse_Scint+Run4D127 step2 max memory diff 45.2 exceeds +/- 30.0 MiB
  • Error: Workflow 37700.0_CloseByPGun_CE_H_Coarse_Scint+Run4D127 step5 max memory diff 44.5 exceeds +/- 30.0 MiB
  • Error: Workflow 37700.0_CloseByPGun_CE_H_Coarse_Scint+Run4D127 step3 max memory diff 45.8 exceeds +/- 30.0 MiB
  • Error: Workflow 37834.0_TTbar_14TeV+Run4D127PU step3 max memory diff 48.1 exceeds +/- 30.0 MiB
  • Error: Workflow 37834.0_TTbar_14TeV+Run4D127PU step2 max memory diff 45.9 exceeds +/- 30.0 MiB

Max Memory Comparisons exceeding threshold NVIDIA_H100

@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold:

Expand to see workflows ...
  • Error: Workflow 37634.402_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka step2 max memory diff 65.7 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.402_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka step3 max memory diff 67.7 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.403_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Validation step3 max memory diff 67.7 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.403_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Validation step2 max memory diff 65.7 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.404_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Profiling step2 max memory diff 65.0 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.404_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Profiling step3 max memory diff 74.7 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.712_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnly step2 max memory diff 64.8 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.712_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnly step3 max memory diff 48.3 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.713_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnlyAlpakaValidationLST step3 max memory diff 48.5 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.713_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnlyAlpakaValidationLST step2 max memory diff 66.4 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.7503_TTbar_14TeV+Run4D127_HLTHeterogeneousValid step2 max memory diff 57.5 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.751_TTbar_14TeV+Run4D127_HLT75e33TimingAlpaka step2 max memory diff 68.2 exceeds +/- 30.0 MiB

Max Memory Comparisons exceeding threshold NVIDIA_L40S

@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold:

Expand to see workflows ...
  • Error: Workflow 37634.402_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka step3 max memory diff 67.2 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.402_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka step2 max memory diff 64.5 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.403_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Validation step3 max memory diff 66.5 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.403_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Validation step2 max memory diff 64.4 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.404_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Profiling step2 max memory diff 64.6 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.404_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Profiling step3 max memory diff 74.3 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.712_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnly step2 max memory diff 64.7 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.712_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnly step3 max memory diff 48.5 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.713_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnlyAlpakaValidationLST step3 max memory diff 49.4 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.713_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnlyAlpakaValidationLST step2 max memory diff 64.7 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.7503_TTbar_14TeV+Run4D127_HLTHeterogeneousValid step2 max memory diff -124.4 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.751_TTbar_14TeV+Run4D127_HLT75e33TimingAlpaka step2 max memory diff 66.2 exceeds +/- 30.0 MiB

Max Memory Comparisons exceeding threshold NVIDIA_T4

@cms-sw/core-l2 , I found 12 workflow step(s) with memory usage exceeding the error threshold:

Expand to see workflows ...
  • Error: Workflow 37634.402_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka step2 max memory diff 64.1 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.402_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka step3 max memory diff 67.4 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.403_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Validation step3 max memory diff 67.3 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.403_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Validation step2 max memory diff 64.0 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.404_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Profiling step2 max memory diff 64.1 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.404_TTbar_14TeV+Run4D127_Patatrack_PixelOnlyAlpaka_Profiling step3 max memory diff 74.4 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.712_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnly step3 max memory diff 49.0 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.712_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnly step2 max memory diff 64.2 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.713_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnlyAlpakaValidationLST step2 max memory diff 64.2 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.713_TTbar_14TeV+Run4D127_lstmkFitOnGPUIters01TrackingOnlyAlpakaValidationLST step3 max memory diff 47.8 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.7503_TTbar_14TeV+Run4D127_HLTHeterogeneousValid step2 max memory diff 115.2 exceeds +/- 30.0 MiB
  • Error: Workflow 37634.751_TTbar_14TeV+Run4D127_HLT75e33TimingAlpaka step2 max memory diff 67.0 exceeds +/- 30.0 MiB

@smuzaffar

Copy link
Copy Markdown
Contributor Author

ignore tests-rejected with external-failure

unit tests failure are not related to this change. failure are already in IBs due to #51649

@smuzaffar

Copy link
Copy Markdown
Contributor Author

@cms-sw/dqm-l2 , @cms-sw/reconstruction-l2 @cms-sw/ml-l2 can you please review this PR. This is needed for Tensorflowversion 2.21.0 along with other externals needed by TF 2.21.0 ( see cms-sw/cmsdist#10789 (comment) for details). Note that there is a increase of 40-60MB for max memory used by relvals. This could easily be due to newer TF and absl version (which has far more shared libs as compare to old absl version : https://github.com/cms-sw/cmsdist/pull/10789/changes#diff-ca8f2ece019f7963a03325e0d321161afe29aecec57373401d2f4e531828a249 ).

Integrating TF 2.21.0 will allow us to move forward with pythn version update too (current TF version doe snot support newer Python versions)

@rseidita

Copy link
Copy Markdown
Contributor

+dqm

@Moanwar

Moanwar commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

+1

@cmsbuild

Copy link
Copy Markdown
Contributor

REMINDER @ftenchini, @mandrenguyen, @sextonkennedy: This PR was tested with cms-sw/cmsdist#10789, please check if they should be merged together

@smuzaffar

Copy link
Copy Markdown
Contributor Author

@cms-sw/orp-l2 can we get this in 20.1.X IBs ? Once you merge it then please also sign cms-sw/cmsdist#10789 . This will allow us to move forward with TF 2.21

@smuzaffar

Copy link
Copy Markdown
Contributor Author

@cms-sw/ml-l2 ( @y19y19 ) can you please review this? This allows us to move forward with TF 2.21 integration in 20.1.X

@smuzaffar

Copy link
Copy Markdown
Contributor Author

FYI @valsdav , @hjkwon260 @y19y19 , can you please review this ?

@valsdav

valsdav commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

+ml

@cmsbuild

Copy link
Copy Markdown
Contributor

This pull request is fully signed and it will be integrated in one of the next master IBs (test failures were overridden). This pull request will now be reviewed by the release team before it's merged. @mandrenguyen, @sextonkennedy, @ftenchini (and backports should be raised in the release meeting by the corresponding L2)
Notice This PR was tested with additional Pull Request(s), please also merge them if necessary: cms-sw/cmsdist#10789

@mandrenguyen

Copy link
Copy Markdown
Contributor

+1

@cmsbuild
cmsbuild merged commit b254ff6 into cms-sw:master Aug 31, 2026
24 of 25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants