Skip to content

Keep audio composition tracks on their own sample clock - #449

Open
Redth wants to merge 1 commit into
sw33tLie:mainfrom
Redth:fix/audio-composition-timescale
Open

Redth wants to merge 1 commit into
sw33tLie:mainfrom
Redth:fix/audio-composition-timescale

Conversation

@Redth

@Redth Redth commented Oct 5, 2026

Copy link
Copy Markdown

Fixes #448.

What's broken

VideoCompositionBuilder.build assigns the inflated editing timescale to the audio composition track:

let timeScale = editingTimeScale(tracks: [sourceVideo] + sourceAudio, frameDuration: frameDuration)
video.naturalTimeScale = timeScale
...
track.naturalTimeScale = timeScale   // audio

editingTimeScale deliberately scales the common clock up toward 1 GHz. For a 48 kHz system-audio take at 30 fps that is LCM(60000, 48000, 30) = 240000, inflated to 240000 * 4166 = 999840000 — exactly the timescale in the bad files.

When a take has a single audio track the export passes that track through instead of mixing it, so the composition track's naturalTimeScale lands in the delivered file as the audio mdhd timescale. ffmpeg-family demuxers then measure the edit list's priming media_time against the sample rate rather than the media timescale:

33,048,132 / 999,840,000 =   0.033 s   <- intended priming skip (~1587 samples)
33,048,132 /      48,000 = 688.50 s    <- what gets applied

688 s of skip against an 11 s track discards every packet. The demuxer reports no error and the decoder emits nothing, so the export is silent on Discord, YouTube and in browsers while AVFoundation still plays it correctly.

Two audio tracks (mic + system) have to be mixed, which re-encodes and resets the clock to 48000 — which is why the existing RecordingMediaFixture.mixedMovie fixture never caught this, and why the bug only shows up for single-track takes.

The fix

Audio keeps the source track's own sample clock. Video is untouched and still gets the high-resolution editing clock it needs for exact endpoints.

track.naturalTimeScale = source.naturalTimeScale > 0 ? source.naturalTimeScale : timeScale

Rounding audio edit boundaries to 1/48000 is the granularity the media actually has, so nothing is lost by not storing them on a 1 GHz clock.

Verification

Measured end to end on a real 64 s system-audio recording, exported through the .high preset path (AVAssetExportSession + videoComposition):

audio mdhd timescale ffmpeg decoded samples
before 999840000 0
after 48000 6133632 (mean −21.9 dB, peak −0.3 dB)
# before
$ ffmpeg -i export.mp4 -map 0:a -af volumedetect -f null -
535 packets read (383242 bytes); 0 frames decoded; 0 decode errors (0 samples)
n_samples: 0

# after
n_samples: 6133632
mean_volume: -21.9 dB
max_volume:  -0.3 dB

The same export also reproduced the original report's exact track layout (soun timescale 999840000 as track 1, vide timescale 600 as track 2), which is what first pointed at this code path rather than MP4WriterSession — the durable original is written correctly at 48000.

Tests

VideoCompositionAudioClockTests covers both the unit invariant and the delivered file. It needs a single-audio-track source, so it builds one with MP4WriterSession (recordSystemAudio: true, recordMicAudio: false) rather than using mixedMovie, whose two tracks mask the bug.

Both tests fail on main with ("999840000") is not equal to ("48000") and pass with this change. The assertions are on the declared clock rather than on decodability, because AVAssetReader happily decodes the broken file — tolerance on Apple's side is the reason this ships looking fine.

  • Builds without warnings
  • scripts/run-tests.sh passes — 820 passed, 0 failed
  • Both new tests verified failing on main and passing here
  • Not exercised through the GUI; verification was the automated export path above plus the standalone harness on a real recording

Not included

Two adjacent observations from the issue, left out to keep this focused — happy to add either here or in a follow-up:

  • The four audio AVAssetWriterInputs (MP4WriterSession ×2, VideoTranscoder, AudioTrackMixer) never set mediaTimeScale, while every video input does. Harmless today since samples arrive at 48 kHz, but it would make the writers robust against an imported asset carrying an odd clock.
  • Recordings are written with moov after mdat, which prevents progressive playback on the web.

🤖 Generated with Claude Code

The editing timescale is inflated toward 1 GHz so edited video endpoints
stay exact, and the audio composition track inherited it. When a take has
a single audio track the export passes that track through rather than
mixing it, so the inflated clock reached the delivered file as the audio
mdhd timescale (999840000 for a 48 kHz take at 30 fps).

ffmpeg-family demuxers then measure the edit list's priming media_time
against the sample rate instead of the media timescale. On an affected
file the intended 33 ms skip is read as 688 s, which overshoots the
track, so every packet is discarded and the export plays silently on
Discord, YouTube and in browsers while AVFoundation still sounds right.

Audio now keeps the source track's sample clock. Video is unchanged and
still gets the high-resolution editing clock. Verified on a 64 s system
audio recording: the exported audio track goes from timescale 999840000
decoding to 0 samples, to timescale 48000 decoding all 6133632 samples.

Fixes sw33tLie#448

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

System Audio track malformed in recording

1 participant