Narration
Recorded or synthesised speech plus a WebVTT timing track that drives captions and word-by-word highlighting.
npx susegad add narration
Stands on: Core, Core: components
The prompt
the prompt
Recorded or synthesised speech, a WebVTT timing track, captions that are always present, and word-by-word highlighting as the audio plays.
<link rel="stylesheet" href="susegad/narration/narration.css">
<script type="module" src="susegad/narration/reader.js"></script>
<sg-narration lang="en-IN">
<audio controls preload="none">
<source src="story.wav" type="audio/wav">
<track kind="captions" src="story.vtt" default>
</audio>
<p><span data-word>Rain</span> <span data-word>came</span> <span data-word>to</span> <span data-word>the</span> <span data-word>window.</span></p>
</sg-narration>
The prompt
Build narration for a document or a story: a phrase splitter that never exceeds a provider's character limit and never breaks a word, treating the Devanagari danda and double danda as sentence enders alongside the Latin ones. Because no text-to-speech API gives word-level timestamps, measure each phrase's real audio duration from its WAV header (never trust a provider's own word count) and spread it over the phrase's words by a script-aware weight: a vowel-group count for Latin, and a proper akshara count for Devanagari and Kannada, where a consonant starts a new akshara unless the one before it ends in a virama (a conjunct stays one akshara), and a matra or nasal attaches to the current akshara rather than starting a new one. Hold a pause after punctuation, longer at a sentence end than a comma. Write the result as WebVTT, one cue per phrase for captions, with a timestamp tag before every word after the first so a highlight driver can mark the current one, in the format the spec already defines so a caption renderer that doesn't understand the tags still shows plain words. Cache every synthesised phrase on disk under a hash of exactly what changes the audio (the text, language, voice, model and pace), so tests and demos never call the network and a repeat build costs nothing; write the request beside the audio, never the key. Build a provider interface any text-to-speech service can implement, a silent-WAV stub that makes every test and demo work offline, and a real adapter, confirmed against the provider's current docs rather than assumed from memory or an old default. Drive highlighting from a native <audio> and a native <track kind="captions">, so captions work with no JavaScript at all and nothing autoplays; with JavaScript, mark the current word with aria-current, matching words in playback order so a repeated word still lines up, and reset that matching on a scrub so seeking never leaves a stale word marked. Fall back to the browser's own voice, or to the page's own text and captions, when no audio can be reached at all.
Words to code
| When you say | Technique | What happens |
|---|---|---|
| never exceeds a provider's character limit and never breaks a word | phrase splitting | splitPhrases(text, { maxChars }) breaks at sentence ends first, then clause punctuation, then word boundaries, joining pieces back up to just under the limit. |
| measure each phrase's real audio duration from its WAV header | honesty | parseWav(bytes) walks the RIFF chunks (fmt and data aren't always adjacent) and divides the data chunk's byte length by the frame size and sample rate — the true duration, not an estimate. |
| a script-aware weight … a proper akshara count | the akshara rules | wordWeight(word) in weights.js: for Devanagari and Kannada, a state machine walks the codepoints, counting a consonant unless the previous character was a virama (which joins it into a conjunct with the consonant before), and letting matras, anusvara and visarga attach without incrementing the count; an independent vowel always counts. |
| hold a pause after punctuation, longer at a sentence end than a comma | pacing | timing.js's PAUSE_AFTER table gives each punctuation mark a pause in word-weight-units (the danda and double danda included), added to the total the phrase's duration is divided by. |
| a highlight driver can mark the current one … a caption renderer that doesn't understand the tags still shows plain words | WebVTT karaoke tags | writeVtt puts a <HH:MM:SS.mmm> tag before every word after a cue's first, exactly the form the WebVTT spec defines for this; parseVtt reads it back the same way. |
| cache every synthesised phrase on disk under a hash | the fixture cache | fixtureKey({ text, lang, voice, model, pace }) hashes only the fields that change the audio; cache.js writes audio.wav and a request.json with no key in it, ever. |
| a provider interface … a real adapter, confirmed against the provider's current docs | ports and adapters | providers/index.js defines the shape; providers/stub.js and providers/sarvam.js both implement synthesize(req); the Sarvam adapter's endpoint, fields and model strings were read from docs.sarvam.ai on the day it was built, recorded in the Wave 4 brief, not assumed. |
drive highlighting from a native <audio> and a native <track kind="captions"> | native first | <sg-narration> enhances the native elements; without JavaScript they already give a person captioned audio with a play button. |
| matching words in playback order so a repeated word still lines up | stateful matching | #mark walks forward from the last matched span's index rather than searching from the start each time, so two occurrences of "the" resolve to the right one as the story plays. |
| reset that matching on a scrub | correctness under seeking | the seeked event clears the match state, so the next tick searches fresh instead of getting stuck past the seek point. |
| fall back to the browser's own voice, or to the page's own text and captions | graceful degradation | speakFallback(text, opts) calls speechSynthesis, reporting which mode ran; when neither that nor real audio is reachable, the static text and native captions already on the page are the narration. |
Accessibility
- Captions are always present; the highlight never carries information the captions don't already give.
- The play control is the consent (decision 0015); narration never autoplays and is never gated by the ambient sound switch.
aria-current="true"is the only signal on a word span — no colour-only cue.- Pronunciation, register and word choice for a real script are a native speaker's job;
story/spread.json'sowner_to_confirmandpronunciation_notesare the pattern for flagging what needs that review before it ships.
Credit
The word-timing method — measure real audio, then spread by a script-aware syllable weight — is a workaround for a real gap: no mainstream Indian-language text-to-speech API publishes word-level timestamps as of this writing. It is an estimate, tested to read naturally, not a claim of precision.
Recorded or synthesised speech, plus a WebVTT timing track that drives captions and word-by-word highlighting. Captions are always available; the play control is the consent (decision 0015) — nothing plays on its own.
The pieces
phrase.js—splitPhrases(text, { maxChars = 220 }). Breaks a script at sentence ends (including the Devanagari danda।and double danda॥), then clause punctuation, then word boundaries, so a call never exceeds Sarvam's 2,500-character limit and never splits a word.wav.js—parseWav(bytes)reads a WAV's real sample rate, channel count, bit depth and duration by walking its chunks (never trusting a provider's own word count);writeSilentWav(opts)makes a deterministic silent WAV for the stub provider;concatWav(parts)lays several clips end to end with silence held between them, for one playable track.weights.js—wordWeight(word): a Latin word's weight is its vowel-group count; a Devanagari or Kannada word's weight is its akshara count, with a real segmenter (a consonant starts a new akshara unless the one before it ends in a virama, which makes a conjunct; matras, anusvara and visarga attach to the current akshara; an independent vowel starts its own).timing.js—timePhrase({ text, durationSec })spreads a phrase's real, WAV-measured duration over its words by weight, holding a pause after punctuation (longer at a sentence end or the danda than at a comma).vtt.js—writeVtt(phrases)/parseVtt(text): one WebVTT cue per phrase, with a<HH:MM:SS.mmm>tag before every word after the first, in the standard karaoke-tag form, so a caption renderer that ignores the tags still shows plain words.cache.js— the fixture cache:fixtureKey({ text, lang, voice, model, pace })hashes exactly the fields that change the audio;readFixture/writeFixturestoreaudio.wavandrequest.json(never a key) underfixtures/<key>/.providers/—synthesize({ text, lang, voice, model, pace }) => { audio, mime, meta }.stub.jsmakes a silent WAV of a plausible length, no network, deterministic.sarvam.jscalls Sarvam's bulbul text-to-speech (POST https://api.sarvam.ai/text-to-speech), reading the key fromSARVAM_API_KEYat call time only.index.js—synthesizeScript(text, opts)andsynthesizePhrases(phraseList, opts)tie it together: split (or take a pre-split script), read the cache, call the provider on a miss, measure the real duration, spread the words, and lay phrases end to end into one track with a matching VTT.synth.mjs— a CLI that fills the fixture cache from a script file, calling the network only on a miss.reader.js—<sg-narration>, the highlight driver, andspeakFallback(text, opts), the offline fallback overspeechSynthesis.
<sg-narration>
<sg-narration lang="en-IN">
<audio controls preload="none">
<source src="story.wav" type="audio/wav">
<track kind="captions" src="story.vtt" default>
</audio>
<p><span data-word>Rain</span> <span data-word>came</span> <span data-word>to</span> <span data-word>the</span> <span data-word>window.</span></p>
</sg-narration>
- The native
<audio>and its<track kind="captions">are the whole thing without JavaScript: captions work, nothing autoplays. - With JavaScript, the element marks the current
[data-word]span witharia-current="true"as the audio plays, matching words in playback order so a repeated word still lines up. No live-region chatter — the words are already visible as captions or as the page's own text. [data-word]spans are optional: build them however fits your page (the storybook-spread recipe'swrapWords(el)wraps a paragraph's text automatically). Without them,<sg-narration>still drives native captions and firessg-word({ word, start, end }) so a page can render its own highlight, andsg-phrase({ index, start, end }) once per phrase, useful for driving a scene's params off the same beats as the narration.- A scrub (forward or back) resets the match so it searches from the start again, so seeking never leaves a stale word marked.
- The track is read independently of the browser's own caption loading (its
loadevent can race the custom element's upgrade), so highlighting works even whenpreload="none"delays the native track fetch.
The offline fallback
speakFallback(text, { lang, onWord, onDone }) speaks with the browser's own voice when no audio is available at all (no key at build time, a network-free demo, or a failed <audio> load). It reports 'speech-synthesis' when it started speaking, or 'captions-only' when speechSynthesis doesn't exist — either way, the page's own text and captions are the narration a person can always fall back to, which the charter requires regardless of which mode ran.
Building a script
import { synthesizePhrases } from 'susegad/narration/index.js';
import sarvam from 'susegad/narration/providers/sarvam.js';
const result = await synthesizePhrases(
[{ beat: 'sky-turns', text: 'The sky turned the colour of an old kadai.', pace: 0.88, pause_after: 'short' }, /* … */],
{ lang: 'en-IN', voice: 'shubh', model: 'bulbul:v3', provider: sarvam },
);
// result.audio: one WAV, phrases laid end to end with held silence between them
// result.vtt: one WebVTT track, captions and word highlighting together
// result.phrases[i].beat: kept from the input, for driving a scene's cues off the same beats
Every call goes through the fixture cache first; a repeat build costs nothing. synthesizeScript(text, opts) does the same from a plain string, splitting it with splitPhrases first.
Accessibility and honesty
- Captions are always present; nothing about the story is said only by the highlight.
- The play control is the only consent narration needs (decision 0015); it is never wired to the global sound switch and never autoplays.
aria-currentis the only state a word span carries — no colour-only signal, and the transition (seenarration.css) is skipped under reduced motion by relying on the CSSprefers-reduced-motion: no-preferenceguard, so word changes still happen, just without the colour transition's motion.- Pronunciation and register are a native speaker's job, not this package's:
story/spread.json'spronunciation_notesandowner_to_confirmare the pattern to follow for a new script.
Known limits
- No word-level timestamps come from Sarvam;
timing.js's weights are an estimate, not a measurement. They read naturally in testing but are not perfectly synced to the real audio's stresses. - Seeking while paused does not recompute the highlighted word (only a scrub during playback does); this is a minor, documented gap, not a broken state — nothing incorrect is shown, only nothing until play resumes.
concatWavrequires every clip to share sample rate, channels and bit depth, which holds for a script synthesised in one pass with one provider, but would need resampling to mix providers within a track.