Multimodal — STT, TTS, vision

Language models eat tokens. Multimodal systems add encoders that map audio or images into that shared space (or into tool side-channels).

Multimodal I/O schematic

Speech → text (STT / ASR)

Automatic speech recognition turns waveforms into text (or token lattices). Modern stacks use encoder-decoders or CTC/transducer hybrids; many products wrap cloud STT behind an API.

Spark: listen companions, dry stubs, gated live STT — VOICE.md. Not production telephony.

Text → speech (TTS)

TTS maps text (or phonemes) to audio. Neural vocoders dominate; latency and voice cloning ethics matter in production.

Spark: speak companions, gated live — same voice doc. Spark train stays on CPU or a consumer GPU.

Vision

Image encoders (ViT-style patches, CNN towers) project pixels into embeddings the LLM can attend to — “see” as tokens, not magic.

Spark: eyes / vision runtime is a stubModel aspects. Do not invent live vision.

flowchart LR
         mic[Audio] --> stt[STT]
         cam[Image] --> vit[Vision encoder]
         stt --> llm[LLM / policy]
         vit --> llm
         llm --> txt[Text]
         llm --> tts[TTS]
  • Ears → brain → voice SVG lives under docs/images/ (model aspects).
  • Workflow: /workflow.html
  • Hive home: Knowledge

Next: Agents & tools · Safety.