Repository navigation
Conversation
The editing timescale is inflated toward 1 GHz so edited video endpoints stay exact, and the audio composition track inherited it. When a take has a single audio track the export passes that track through rather than mixing it, so the inflated clock reached the delivered file as the audio mdhd timescale (999840000 for a 48 kHz take at 30 fps). ffmpeg-family demuxers then measure the edit list's priming media_time against the sample rate instead of the media timescale. On an affected file the intended 33 ms skip is read as 688 s, which overshoots the track, so every packet is discarded and the export plays silently on Discord, YouTube and in browsers while AVFoundation still sounds right. Audio now keeps the source track's sample clock. Video is unchanged and still gets the high-resolution editing clock. Verified on a 64 s system audio recording: the exported audio track goes from timescale 999840000 decoding to 0 samples, to timescale 48000 decoding all 6133632 samples. Fixes sw33tLie#448 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #448.
What's broken
VideoCompositionBuilder.buildassigns the inflated editing timescale to the audio composition track:editingTimeScaledeliberately scales the common clock up toward 1 GHz. For a 48 kHz system-audio take at 30 fps that isLCM(60000, 48000, 30) = 240000, inflated to240000 * 4166 =999840000— exactly the timescale in the bad files.When a take has a single audio track the export passes that track through instead of mixing it, so the composition track's
naturalTimeScalelands in the delivered file as the audiomdhdtimescale. ffmpeg-family demuxers then measure the edit list's primingmedia_timeagainst the sample rate rather than the media timescale:688 s of skip against an 11 s track discards every packet. The demuxer reports no error and the decoder emits nothing, so the export is silent on Discord, YouTube and in browsers while AVFoundation still plays it correctly.
Two audio tracks (mic + system) have to be mixed, which re-encodes and resets the clock to 48000 — which is why the existing
RecordingMediaFixture.mixedMoviefixture never caught this, and why the bug only shows up for single-track takes.The fix
Audio keeps the source track's own sample clock. Video is untouched and still gets the high-resolution editing clock it needs for exact endpoints.
Rounding audio edit boundaries to 1/48000 is the granularity the media actually has, so nothing is lost by not storing them on a 1 GHz clock.
Verification
Measured end to end on a real 64 s system-audio recording, exported through the
.highpreset path (AVAssetExportSession+videoComposition):mdhdtimescale99984000048000The same export also reproduced the original report's exact track layout (
sountimescale 999840000 as track 1,videtimescale 600 as track 2), which is what first pointed at this code path rather thanMP4WriterSession— the durable original is written correctly at 48000.Tests
VideoCompositionAudioClockTestscovers both the unit invariant and the delivered file. It needs a single-audio-track source, so it builds one withMP4WriterSession(recordSystemAudio: true, recordMicAudio: false) rather than usingmixedMovie, whose two tracks mask the bug.Both tests fail on
mainwith("999840000") is not equal to ("48000")and pass with this change. The assertions are on the declared clock rather than on decodability, becauseAVAssetReaderhappily decodes the broken file — tolerance on Apple's side is the reason this ships looking fine.scripts/run-tests.shpasses — 820 passed, 0 failedmainand passing hereNot included
Two adjacent observations from the issue, left out to keep this focused — happy to add either here or in a follow-up:
AVAssetWriterInputs (MP4WriterSession×2,VideoTranscoder,AudioTrackMixer) never setmediaTimeScale, while every video input does. Harmless today since samples arrive at 48 kHz, but it would make the writers robust against an imported asset carrying an odd clock.moovaftermdat, which prevents progressive playback on the web.🤖 Generated with Claude Code