Skip to content

Text to Speech page: Speak Text section with voice menu and markup guide - #64

Merged
dsward2 merged 1 commit into
mainfrom
feature/tts-speak-text
Sep 29, 2026
Merged

dsward2 merged 1 commit into
mainfrom
feature/tts-speak-text

Conversation

@dsward2

@dsward2 dsward2 commented Sep 29, 2026

Copy link
Copy Markdown
Owner

What

A new Speak Text section at the bottom of the web UI's Text to Speech page (Devices):

  • A text box, prefilled with <speak>It is four oh nine <break time="700ms"/> <prosody rate="80%">on AntennaHead Radio.</prosody></speak>.
  • A Voice menu of the installed voices. "Default voice" means the Text to Speech voice chosen in Configuration.
  • A Speak button that speaks the text once through the live pipeline (PCMSpeechSynth → sox → PCMUDPSender) and starts the web player, the same way Listen does. When it finishes, the filler starts, as it does after a one-pass Listen.
  • Below that, a collapsible guide, "Controlling the voice: pauses, speed, pronunciation", covering the two markup methods:
    • SSML for modern voices, with a table of which tags actually changed the audio when I rendered them with the Premium Ava voice. prosody rate and volume, break, say-as characters and phoneme work; prosody pitch barely changes anything; emphasis and sub are ignored.
    • [[slnc]], [[rate]], [[pbas]] and [[pmod]] embedded commands for classic voices.
    • A warning not to mix the two.

How

  • Text that starts with <speak is sent with --ssml. That's the same rule StationDirector's announcer uses, and the rule is needed. In tests with file input:
    • SSML with --ssml: 3.84 s, with the pause and the slower ending.
    • Plain text with --ssml: no audio at all.
    • SSML without --ssml: 10.8 s, because the tags are read aloud.
  • SDRController: the launch half of startTextToSpeech moved into startSpeechSynthPipeline(text:repeatForever:voiceIdentifier:ssml:dying:), which the new speakText(_:voiceIdentifier:) shares. makeSpeechSynthTaskItem accepts an optional voice and SSML flag. Existing folder behavior is unchanged: same voice setting, no --ssml.
  • New route: POST /speaktextbuttonclicked.html with body {text, voice}.
  • antennahead.js: speakTextButtonClicked(form).

Verification

  • Debug build succeeds, and node --check passes on antennahead.js.
  • I tested PCMSpeechSynth's file input with and without --ssml as described above.
  • Not yet tested in the running app. That needs the installed AntennaHead restarted, which interrupts the live stream.

🤖 Generated with Claude Code

A new section at the bottom of the web UI's Text to Speech page: a text
box (prefilled with an SSML example), a Voice menu of installed voices
("Default voice" = the Text to Speech voice setting), and a Speak button
that speaks the text once through the live pipeline. Text starting with
<speak is passed to PCMSpeechSynth with --ssml, the same rule
StationDirector uses; plain text with --ssml renders nothing, and SSML
without it has its tags read aloud.

Below it, a collapsible guide to the two markup methods: SSML for modern
voices (with what was measured to work with the Premium Ava voice) and
[[...]] embedded commands for classic voices.

SDRController: the launch half of startTextToSpeech moves into
startSpeechSynthPipeline, shared with the new speakText(_:voiceIdentifier:).
Route: POST /speaktextbuttonclicked.html {text, voice}.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dsward2
dsward2 merged commit a309fa1 into main Sep 29, 2026
1 check passed
@dsward2
dsward2 deleted the feature/tts-speak-text branch September 29, 2026 00:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant