fix(model): use dynamic temporal video token counts - #5590
Draft
cuichenx wants to merge 1 commit into
Draft
Conversation
Signed-off-by: Chen Cui <chcui@nvidia.com>
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do ?
Adds an explicit, processor-driven temporal-video resize mode for Nemotron Omni Energon training so each tubelet receives the number of image placeholders its actual post-resize vision grid produces.
Changelog
fixed_512as the compatibility default and enableprocessormode only in the shipped Nemotron Omni VALOR Energon recipe.imgs_sizes, and reported token counts instead of reimplementing its resize policy.num_image_tilesmetadata separate.GitHub Actions CI
Focused validation on one H100:
uv run --no-sync pre-commit run --all-files384x672frames, five 252-token tubelets, and 1,260 placeholders matching 1,260 expected vision rows.git diff --checkBefore your PR is "Ready for review"
Pre checks:
Additional Information
This PR intentionally scopes dynamic temporal-video sizing to the canonical Energon path. The fixed mode remains available for compatibility, and the deprecated LLaVA contract remains fixed-resolution.