Susegad UI
Register
Theme
Palette

Foundations

Narration

Recorded or synthesised speech plus a WebVTT timing track that drives captions and word-by-word highlighting.

npx susegad add narration

Stands on: Core, Core: components

The prompt

the prompt

Recorded or synthesised speech, a WebVTT timing track, captions that are always present, and word-by-word highlighting as the audio plays.

<link rel="stylesheet" href="susegad/narration/narration.css">
<script type="module" src="susegad/narration/reader.js"></script>

<sg-narration lang="en-IN">
  <audio controls preload="none">
    <source src="story.wav" type="audio/wav">
    <track kind="captions" src="story.vtt" default>
  </audio>
  <p><span data-word>Rain</span> <span data-word>came</span> <span data-word>to</span> <span data-word>the</span> <span data-word>window.</span></p>
</sg-narration>

The prompt

Build narration for a document or a story: a phrase splitter that never exceeds a provider's character limit and never breaks a word, treating the Devanagari danda and double danda as sentence enders alongside the Latin ones. Because no text-to-speech API gives word-level timestamps, measure each phrase's real audio duration from its WAV header (never trust a provider's own word count) and spread it over the phrase's words by a script-aware weight: a vowel-group count for Latin, and a proper akshara count for Devanagari and Kannada, where a consonant starts a new akshara unless the one before it ends in a virama (a conjunct stays one akshara), and a matra or nasal attaches to the current akshara rather than starting a new one. Hold a pause after punctuation, longer at a sentence end than a comma. Write the result as WebVTT, one cue per phrase for captions, with a timestamp tag before every word after the first so a highlight driver can mark the current one, in the format the spec already defines so a caption renderer that doesn't understand the tags still shows plain words. Cache every synthesised phrase on disk under a hash of exactly what changes the audio (the text, language, voice, model and pace), so tests and demos never call the network and a repeat build costs nothing; write the request beside the audio, never the key. Build a provider interface any text-to-speech service can implement, a silent-WAV stub that makes every test and demo work offline, and a real adapter, confirmed against the provider's current docs rather than assumed from memory or an old default. Drive highlighting from a native <audio> and a native <track kind="captions">, so captions work with no JavaScript at all and nothing autoplays; with JavaScript, mark the current word with aria-current, matching words in playback order so a repeated word still lines up, and reset that matching on a scrub so seeking never leaves a stale word marked. Fall back to the browser's own voice, or to the page's own text and captions, when no audio can be reached at all.

Words to code

When you sayTechniqueWhat happens
never exceeds a provider's character limit and never breaks a wordphrase splittingsplitPhrases(text, { maxChars }) breaks at sentence ends first, then clause punctuation, then word boundaries, joining pieces back up to just under the limit.
measure each phrase's real audio duration from its WAV headerhonestyparseWav(bytes) walks the RIFF chunks (fmt and data aren't always adjacent) and divides the data chunk's byte length by the frame size and sample rate — the true duration, not an estimate.
a script-aware weight … a proper akshara countthe akshara ruleswordWeight(word) in weights.js: for Devanagari and Kannada, a state machine walks the codepoints, counting a consonant unless the previous character was a virama (which joins it into a conjunct with the consonant before), and letting matras, anusvara and visarga attach without incrementing the count; an independent vowel always counts.
hold a pause after punctuation, longer at a sentence end than a commapacingtiming.js's PAUSE_AFTER table gives each punctuation mark a pause in word-weight-units (the danda and double danda included), added to the total the phrase's duration is divided by.
a highlight driver can mark the current one … a caption renderer that doesn't understand the tags still shows plain wordsWebVTT karaoke tagswriteVtt puts a <HH:MM:SS.mmm> tag before every word after a cue's first, exactly the form the WebVTT spec defines for this; parseVtt reads it back the same way.
cache every synthesised phrase on disk under a hashthe fixture cachefixtureKey({ text, lang, voice, model, pace }) hashes only the fields that change the audio; cache.js writes audio.wav and a request.json with no key in it, ever.
a provider interface … a real adapter, confirmed against the provider's current docsports and adaptersproviders/index.js defines the shape; providers/stub.js and providers/sarvam.js both implement synthesize(req); the Sarvam adapter's endpoint, fields and model strings were read from docs.sarvam.ai on the day it was built, recorded in the Wave 4 brief, not assumed.
drive highlighting from a native <audio> and a native <track kind="captions">native first<sg-narration> enhances the native elements; without JavaScript they already give a person captioned audio with a play button.
matching words in playback order so a repeated word still lines upstateful matching#mark walks forward from the last matched span's index rather than searching from the start each time, so two occurrences of "the" resolve to the right one as the story plays.
reset that matching on a scrubcorrectness under seekingthe seeked event clears the match state, so the next tick searches fresh instead of getting stuck past the seek point.
fall back to the browser's own voice, or to the page's own text and captionsgraceful degradationspeakFallback(text, opts) calls speechSynthesis, reporting which mode ran; when neither that nor real audio is reachable, the static text and native captions already on the page are the narration.

Accessibility

  • Captions are always present; the highlight never carries information the captions don't already give.
  • The play control is the consent (decision 0015); narration never autoplays and is never gated by the ambient sound switch.
  • aria-current="true" is the only signal on a word span — no colour-only cue.
  • Pronunciation, register and word choice for a real script are a native speaker's job; story/spread.json's owner_to_confirm and pronunciation_notes are the pattern for flagging what needs that review before it ships.

Credit

The word-timing method — measure real audio, then spread by a script-aware syllable weight — is a workaround for a real gap: no mainstream Indian-language text-to-speech API publishes word-level timestamps as of this writing. It is an estimate, tested to read naturally, not a claim of precision.

Recorded or synthesised speech, plus a WebVTT timing track that drives captions and word-by-word highlighting. Captions are always available; the play control is the consent (decision 0015) — nothing plays on its own.

The pieces

  • phrase.js — splitPhrases(text, { maxChars = 220 }). Breaks a script at sentence ends (including the Devanagari danda । and double danda ॥), then clause punctuation, then word boundaries, so a call never exceeds Sarvam's 2,500-character limit and never splits a word.
  • wav.js — parseWav(bytes) reads a WAV's real sample rate, channel count, bit depth and duration by walking its chunks (never trusting a provider's own word count); writeSilentWav(opts) makes a deterministic silent WAV for the stub provider; concatWav(parts) lays several clips end to end with silence held between them, for one playable track.
  • weights.js — wordWeight(word): a Latin word's weight is its vowel-group count; a Devanagari or Kannada word's weight is its akshara count, with a real segmenter (a consonant starts a new akshara unless the one before it ends in a virama, which makes a conjunct; matras, anusvara and visarga attach to the current akshara; an independent vowel starts its own).
  • timing.js — timePhrase({ text, durationSec }) spreads a phrase's real, WAV-measured duration over its words by weight, holding a pause after punctuation (longer at a sentence end or the danda than at a comma).
  • vtt.js — writeVtt(phrases) / parseVtt(text): one WebVTT cue per phrase, with a <HH:MM:SS.mmm> tag before every word after the first, in the standard karaoke-tag form, so a caption renderer that ignores the tags still shows plain words.
  • cache.js — the fixture cache: fixtureKey({ text, lang, voice, model, pace }) hashes exactly the fields that change the audio; readFixture/writeFixture store audio.wav and request.json (never a key) under fixtures/<key>/.
  • providers/ — synthesize({ text, lang, voice, model, pace }) => { audio, mime, meta }. stub.js makes a silent WAV of a plausible length, no network, deterministic. sarvam.js calls Sarvam's bulbul text-to-speech (POST https://api.sarvam.ai/text-to-speech), reading the key from SARVAM_API_KEY at call time only.
  • index.js — synthesizeScript(text, opts) and synthesizePhrases(phraseList, opts) tie it together: split (or take a pre-split script), read the cache, call the provider on a miss, measure the real duration, spread the words, and lay phrases end to end into one track with a matching VTT.
  • synth.mjs — a CLI that fills the fixture cache from a script file, calling the network only on a miss.
  • reader.js — <sg-narration>, the highlight driver, and speakFallback(text, opts), the offline fallback over speechSynthesis.

<sg-narration>

<sg-narration lang="en-IN">
  <audio controls preload="none">
    <source src="story.wav" type="audio/wav">
    <track kind="captions" src="story.vtt" default>
  </audio>
  <p><span data-word>Rain</span> <span data-word>came</span> <span data-word>to</span> <span data-word>the</span> <span data-word>window.</span></p>
</sg-narration>
  • The native <audio> and its <track kind="captions"> are the whole thing without JavaScript: captions work, nothing autoplays.
  • With JavaScript, the element marks the current [data-word] span with aria-current="true" as the audio plays, matching words in playback order so a repeated word still lines up. No live-region chatter — the words are already visible as captions or as the page's own text.
  • [data-word] spans are optional: build them however fits your page (the storybook-spread recipe's wrapWords(el) wraps a paragraph's text automatically). Without them, <sg-narration> still drives native captions and fires sg-word ({ word, start, end }) so a page can render its own highlight, and sg-phrase ({ index, start, end }) once per phrase, useful for driving a scene's params off the same beats as the narration.
  • A scrub (forward or back) resets the match so it searches from the start again, so seeking never leaves a stale word marked.
  • The track is read independently of the browser's own caption loading (its load event can race the custom element's upgrade), so highlighting works even when preload="none" delays the native track fetch.

The offline fallback

speakFallback(text, { lang, onWord, onDone }) speaks with the browser's own voice when no audio is available at all (no key at build time, a network-free demo, or a failed <audio> load). It reports 'speech-synthesis' when it started speaking, or 'captions-only' when speechSynthesis doesn't exist — either way, the page's own text and captions are the narration a person can always fall back to, which the charter requires regardless of which mode ran.

Building a script

import { synthesizePhrases } from 'susegad/narration/index.js';
import sarvam from 'susegad/narration/providers/sarvam.js';

const result = await synthesizePhrases(
  [{ beat: 'sky-turns', text: 'The sky turned the colour of an old kadai.', pace: 0.88, pause_after: 'short' }, /* … */],
  { lang: 'en-IN', voice: 'shubh', model: 'bulbul:v3', provider: sarvam },
);
// result.audio: one WAV, phrases laid end to end with held silence between them
// result.vtt: one WebVTT track, captions and word highlighting together
// result.phrases[i].beat: kept from the input, for driving a scene's cues off the same beats

Every call goes through the fixture cache first; a repeat build costs nothing. synthesizeScript(text, opts) does the same from a plain string, splitting it with splitPhrases first.

Accessibility and honesty

  • Captions are always present; nothing about the story is said only by the highlight.
  • The play control is the only consent narration needs (decision 0015); it is never wired to the global sound switch and never autoplays.
  • aria-current is the only state a word span carries — no colour-only signal, and the transition (see narration.css) is skipped under reduced motion by relying on the CSS prefers-reduced-motion: no-preference guard, so word changes still happen, just without the colour transition's motion.
  • Pronunciation and register are a native speaker's job, not this package's: story/spread.json's pronunciation_notes and owner_to_confirm are the pattern to follow for a new script.

Known limits

  • No word-level timestamps come from Sarvam; timing.js's weights are an estimate, not a measurement. They read naturally in testing but are not perfectly synced to the real audio's stresses.
  • Seeking while paused does not recompute the highlighted word (only a scrub during playback does); this is a minor, documented gap, not a broken state — nothing incorrect is shown, only nothing until play resumes.
  • concatWav requires every clip to share sample rate, channels and bit depth, which holds for a script synthesised in one pass with one provider, but would need resampling to mix providers within a track.