Which voice speaks a line, which voices there are, and the cache of lines already spoken.

A cached line is keyed by the voice and the exact text. What is kept is the audio and the time of every character, not the text: a reader of the store sees spans, and the line is put back from the request that asked for it. Only a line the caller declared fixed is ever kept (see [SpeechCache::keep]).

7use std::fmt;
9use serde::Serialize;
10
11use crate::error::{Diagnostic, StoreError};
12use crate::time::Millis;
13use crate::timing::TimedAudio;
14use crate::voice::SpeechLine;
15
16pub use whiskers_core::{NotAVoiceId, VOICE_ID_MAX, VoiceId};

A voice the account can use: its id and the name a child is told.

19#[derive(Clone, Debug, PartialEq, Eq, Serialize)]
20pub struct VoiceInfo {
21    pub id: VoiceId,
22    pub name: String,
23}

Why the list of voices could not be read.

26#[derive(Clone, Debug, PartialEq, Eq)]
27pub enum VoicesError {
28    NotConfigured,
29    Unreachable(Diagnostic),
30    Refused { status: u16, said: Diagnostic },
31    Unreadable(Diagnostic),
32}
34impl fmt::Display for VoicesError {
35    fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
36        match self {
37            VoicesError::NotConfigured => f.write_str("the voice is not configured"),
38            VoicesError::Unreachable(d) | VoicesError::Unreadable(d) | VoicesError::Refused { said: d, .. } => d.fmt(f),
39        }
40    }
41}
42
43impl std::error::Error for VoicesError {}

What the cache holds.

46#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)]
47pub struct CacheSize {
48    pub entries: u64,

The bytes of audio and timing kept.

50    pub bytes: u64,
51}

Spoken lines kept so that a fixed line costs voice characters once.

Contract, for every adapter:

  • Exact. A line is found only for the voice and the exact text it was kept under.
  • Bounded, least recently used out. Over its cap the adapter drops the entries used longest ago, never the one just kept; a line that alone is over the cap is not kept. find counts as a use.
  • No words. What is kept is audio and spans. The text is the caller's.
  • Only what the caller declared fixed. The service calls keep only for a line its caller said was the same every time (a page name, a hint, a greeting, a preview). Words that came from a memory are never declared, so a forgotten memory leaves nothing here to forget.
64#[expect(async_fn_in_trait, reason = "a Worker's futures hold JavaScript values and cannot be Send, so no Send bound may be required here")]
65pub trait SpeechCache {

The line as kept for this voice, or None.

67    async fn find(&self, voice: &VoiceId, line: &SpeechLine, now: Millis) -> Result<Option<TimedAudio>, StoreError>;

Keeps the line's speech under (voice, text), replacing an older copy.

70    async fn keep(&self, voice: &VoiceId, line: &SpeechLine, speech: &TimedAudio, now: Millis) -> Result<(), StoreError>;
72    async fn size(&self) -> Result<CacheSize, StoreError>;
73}

The bytes one cached line takes against the cap: its audio and eight bytes of timing per character.

76pub fn weight(speech: &TimedAudio) -> u64 {
77    speech.audio().0.len() as u64 + 8 * speech.timing().spans().len() as u64
78}

The cap of the service's speech cache, in bytes: about two thousand lines of speech.

81pub const SPEECH_CACHE_BYTES: u64 = 64 * 1024 * 1024;

The cap of a device's speech cache, in bytes.

84pub const DEVICE_SPEECH_CACHE_BYTES: u64 = 32 * 1024 * 1024;