Voice and sound
AI voices and sound for Foundry VTT
Type one message. Familiar speaks it in the voices you cast, cues the sound, and moves the music with the scene. Your players hear all of it and see the line on the token.
You keep three switches, one each for voice, music and sound effects: run on its own, wait until asked, or stay quiet. It all runs in your browser, on provider keys you bring.
Demo clip coming
What the clip will show
- One message in the chat: the party pushes through the tavern door onto the rain-soaked forest road, and Bram the innkeeper calls after them.
- The narration is spoken in the narrator's voice; Bram answers in his own.
- The door creaks: a sound effect from your own effects playlist, or generated when nothing there matches.
- The tavern music fades to forest ambience from your own playlists, and the mood is bound to the scene.
- On the player's screen the token shows the line marked AI Voice, the line sits in their chat as an AI Voice card, and the player hears everything without clicking anything.
Hear it
Every clip in these bands will be a real take from a live session, recorded with sound, with the ask you read beside it sent word for word. Until a take is recorded its card says what it will prove; a recorded card plays with sound when you press it, never on its own. How each of these works is in the chapters below, starting with speaking the scene.
Same line, six voices
One NPC line, sent through each of the six voice providers with only the provider changed in the Voice tab. Each card will name the model that spoke it and, where the provider reports one, what the line cost.
"Keep your hands where I can see them, and nobody gets hurt."
Demo clip coming
Demo clip coming
Demo clip coming
Demo clip coming
Demo clip coming
Demo clip coming
Same scene, three switches
The message from the top of the page, sent twice. First with voice, ambience and sound effects on Automatic, so one reply speaks, switches the music and plays the creak. Then with all three on On request, and the same reply is text alone.
"The party pushes through the tavern door onto the rain-soaked forest road; Bram the innkeeper calls after them."
Demo clip coming
Demo clip coming
Library first, then generated
Four asks that show where a sound comes from: your own track when the library has one, a generated clip when it does not, a loop filed as ambience, and a positional sound placed on the map.
Demo clip coming
Demo clip coming
Demo clip coming
Demo clip coming
One line, four readings
The same line delivered angry, whispering, sung, and with a pause before it, all on ElevenLabs eleven_v3, the one model that renders every one of these. Other providers take an emotion hint as an instruction, and none of them sings.
"You should not have come back here."
Demo clip coming
Demo clip coming
Demo clip coming
Demo clip coming
The scene, spoken as it is written
Familiar voices the lines it writes, in the same reply, one take per turn.
- With Speak aloud on Automatic, the default, the AI voices the scene lines of its own reply: narration in the narrator's voice, each character in theirs, up to twenty lines in one gapless take. Rules answers, prep and lists stay silent.
- Set it to On request and it waits until you ask.
- Two slots. A narrator voice carries description and the beats between people; a character voice carries the people, and a voice pinned to an NPC wins. The AI picks per line.
- Emotion, a pause before a line, speed and volume are per-provider knobs. Emotion on ElevenLabs needs the eleven_v3 model; speed works everywhere except ElevenLabs; volume is Cartesia only. The table further down says which is which.
- A narration over roughly 280 characters starts speaking at its first sentence boundary while the rest still renders, so a slow voice model gives you the opening line in seconds rather than a long silence. That is Progressive, in the Speech section, on by default while Playback is Immediate.
- Set Playback to Wait for my go and a finished clip holds instead of playing: a note that it is ready, a Play bar in the chat, a pulse on the Familiar button. Press Play, or a key you bind under Configure Controls, or tell your AI to play it, and with broadcast on it starts for you and your players at the same moment, because their copy went ahead while it waited. A held clip expires after ten minutes; Stop drops it; a table chat answer is never held.
- Text over the provider's own cap is split at sentence boundaries and joined into one clip. A line you have heard before comes from the browser cache, for free.
Voice starts switched on, and nothing plays or bills until the Voice tab of the browser that should speak holds a key. Speak aloud is the GM's chat window only; Claude, Codex and the other MCP clients keep calling the voice tools. Automatic spends at most three takes a turn; a fourth is skipped and the reply carries on as text. Progressive applies to Immediate playback after a reply has landed: a voice generated mid-reply, and any held clip, is staged whole.

Demo clip coming
What the clip will show
- One GM message sets a tavern scene with three speakers.
- The reply lands as text and the voices follow in the same turn: narrator, innkeeper, a second NPC, no gap between the lines. The opening line is already playing while the rest is still rendering.
- Each spoken line has its own cell in the chat with the AI Voice badge and the speaker's name, and the whole take lands in the Foundry chat as one AI Voice card with a line per speaker.
- Playback switched to Wait for my go for the last message, broadcast on: the Play bar appears, the Familiar button pulses, and Play starts the line for GM and players together.
Cast a voice without opening a catalogue
Describe the character. The AI finds a voice, casts it on the actor, and tells you which one it picked.
- Ask for a gruff older Scotsman and the AI searches the provider's catalogue by gender, language and free text, casts the best match, and names its choice in the reply.
- A Cast chip in the chat carries a Preview button. On ElevenLabs every stock voice has a free sample; on Cartesia about a quarter do; elsewhere a preview is generated on your key, and the tooltip says which before you click.
- Find a voice yourself in the Voice tab: a search box with gender and language filters and a live count under each voice picker, and a play button beside the picker that auditions the pick, on the narrator slot, the character slot, and Voice Assignments.
- Pin a voice to an actor once and that NPC keeps it every session. An NPC that speaks in table chat without a voice gets one cast on its first spoken reply, and you are whispered which.
- Ask for a different voice model on one take, ElevenLabs v3 for the line that sings, and that take renders on it while the model on your slots stays what it was; the result names the model that spoke.
Labels to search on come from ElevenLabs and Cartesia. OpenAI has thirteen fixed voices, and OpenRouter, NanoGPT and a custom server take a voice name you type. There is no voice picker on the character sheet: casting happens in the chat or in the settings.

Demo clip coming
What the clip will show
- You type: the blacksmith is a gruff older Scotsman, find him a voice.
- A Cast chip appears naming the voice it picked; Preview plays the provider's own sample, free.
- The Voice tab: the Find a Voice bar narrows the catalogue by gender and language, and the play button beside the Default Voice picker above it auditions the pick.
- A second NPC with no voice of its own answers in table chat, gets a voice cast on that first line, and the GM is whispered which one.
Players talk to an NPC in the Foundry chat, and hear it answer
A player types to an NPC in the ordinary Foundry chat. The answer comes back in that NPC's voice.
- A player types @Bram hello in the Foundry chat. Bram answers in his cast voice, spoken from the GM's side, with no tool call and no extra model round. Every NPC a player does not own answers to its name this way out of the box; pin one on its sheet to keep it answering when you switch Any NPC off, in the Chat tab, for a session where a name alone would give something away.
- @familiar reaches the narrator, who answers in the narrator's voice while Speak aloud is on Automatic.
- Mute one NPC from its sheet header, or ask the AI to do it. The AI reads and changes most of Familiar's own settings on request, the sound rows included; a setting that spends your key or reaches the players is the GM's call to widen, and a player's turn can only narrow it.
- A player's own words are never spoken and never become a prompt for a sound. On a player's turn, any tool that changes the world waits for you to confirm.
- Eight seconds between spoken replies from the same NPC. A reply inside that window comes as text.
On request does not silence the table. An NPC in table chat speaks for as long as the Voice tab's own Table chat switch is on, which it is out of the box, and switching it off keeps every table-chat reply as text; only the narrator follows Speak aloud.
Demo clip coming
What the clip will show
- A player's window: they type @Bram hello in the Foundry chat.
- The GM's window posts Bram's reply; the player hears Bram and sees the line in a bubble on his token.
- A second @Bram inside eight seconds comes back as text: the cooldown between spoken replies from one NPC.
Your players hear it without installing anything
Players need no module and no key. Audio reaches them through Foundry's own player, and the line shows on the token.
- Broadcast is Automatic out of the box: every voice clip reaches every connected player as it plays. Manual puts a broadcast button on each clip so you choose; Off keeps voice on your speakers.
- A player hears the clip under their own Foundry interface slider and sees the line in a bubble on the speaking token, marked AI Voice. The line also lands in their Foundry chat as a card marked AI Voice the moment the clip is on its way to them, and a generated one-shot as a card marked AI Sound, so the table can read back what it heard.
- Their music dips to thirty percent under a voice and swells back within a second. That dip is each listener's own: you set yours in the Speech section of the Voice tab, and a player sets theirs on the same two rows in Foundry's Module Settings.
- You keep a Voice volume of your own, a Stop button in the chat header, and your keyboard media keys. Stop All Familiar Audio under Configure Controls, unbound until you pick a chord, silences every clip, drops a held narration and releases the duck.
- Broadcast clips are stored under your world's familiar-tts folder, one file per unique clip, so the same line in the same voice is never uploaded twice.
A player can hear a clip a moment before you do, because the text lands first on your side. A held clip starts for your players at the same moment only when Broadcast was on while it waited; switch it on later and the clip stays with you.
Demo clip coming
What the clip will show
- The player's screen only: no chat window, no settings.
- A voiced line plays; the bubble appears on the token with the AI Voice mark, the card lands in their chat, and the player's music dips and comes back.
- The GM presses Stop mid-line: silence, and the music swells back.
Music and ambience from your own playlists
Say the party pushes into the taproom and Familiar finds the mood in your own playlists, fades over, and binds it to the scene.
- It searches your own Foundry playlists by playlist, folder and track name, with a synonym table that maps tavern, inn and taproom onto the same mood. It never fetches a track and never makes music of its own: it cues what your world already holds, plus any loop you asked Familiar to generate.
- A switch crossfades over two seconds and binds the mood to the active scene, so Foundry cues it again on its own next time.
- Mood switching on Automatic, the default, switches when the scene, the location or a fight changes: one switch a minute, one per turn, and a second ask inside that window comes back as skipped. On request waits for you. Off refuses.
- A Stop all audio button sits on the ambience cell, and the classic playlist tools stay: play, stop, fade, create, reorder, volume, one stop for everything.
- Three ways to get audio in. A pack from the Foundry package browser, where Michael Ghelfi Studios ships one whose playlists arrive as compendium packs you drag into the sidebar. Your own files, through Bulk Import on a playlist's configuration sheet. And services with their own player, which Familiar cannot cue.
Familiar ships no sounds of its own and makes no music. Syrinscape, Tabletop Audio, Kenku FM, YouTube and Spotify play outside Foundry's player, so the AI will not switch them. An empty world gets a hint that points here.

Demo clip coming
What the clip will show
- You type: we push into the taproom.
- The ambience cell lists the tavern moods it found in your playlists; one fades in, and a chip reads bound to Scene.
- Combat starts and the mood switches by itself; a second switch inside the minute is skipped.
Sound effects: your library first, then generated
The door creaks because the narration said so. If your effects playlist has a creak it plays for free; if not, Familiar generates one.
- A track in your own playlist named Sound Effects, SFX, Effects, One-shots or Stingers plays first, at no cost and with no badge.
- When nothing matches, a one-shot is generated on your ElevenLabs voice key or your fal.ai image key: one to six seconds on Automatic, up to twenty-two on request, with the provider's own price printed on the cell where it reports one.
- Automatic means at most two effects a reply and three generated clips a minute; the next one waits. On request plays only when you ask. Off refuses.
- Save keeps a clip as a track in a Familiar Sound Effects playlist. Ask for a rain loop and it is filed in Familiar Ambience, where the music layer finds it. Ask for a waterfall at the cliff and it becomes a positional sound on the map that grows louder as a token walks up.
- Players hear effects through Foundry's own audio while Players hear sound effects is on, the default, and switching it off keeps effects on your own speakers; generated clips are then saved under your world's familiar-sfx folder.
A generated effect bills whichever key you have, so an image key alone is enough to be charged for sound effects, and Sound effects starts on Automatic; set it to On request or Off to hold it. Never music, and loops only when you ask. Effects never duck the music and never wait for a voice line.
Demo clip coming
What the clip will show
- The narration says the old door creaks; your own creak track plays, no badge.
- Thunder rolls with no track to match: a clip is generated, the cell carries the AI Sound badge, and Save files it.
- You ask for a rain loop; it lands in Familiar Ambience, and the next scene switch can cue it.
The other direction: the table, written down
Switch on transcription. The session is written down as it is spoken, into a journal you can search. Next week the AI reads it back to you.
- Your microphone streams to Gladia, Deepgram or AssemblyAI on your key, and the words land in a panel in the chat window. The audio itself is never stored, only the text.
- Deepgram and AssemblyAI label the speakers. Rename Speaker 0 to Ben once and every line, past and future, follows, across sessions.
- Search the transcript with Ctrl+F, drop a bookmark with F9, export in five formats including SRT. The journal fills every thirty seconds.
- After the session, three settings: the transcript becomes searchable for the AI (on by default), so what did we decide about the duke gets an answer next week; a recap page written on your chat model (opt-in); and players read along live in a read-only panel (opt-in).
- Dictation is the same stack pointed at the chat box. Hold Ctrl+Shift+D and speak, or go hands-free and let a pause send.
Gladia, the default, does not tell speakers apart live. The microphone needs a secure page: https, or Foundry opened on localhost. One key does not cover both speech and voices; transcription has its own tab and its own key.

Demo clip coming
What the clip will show
- The mic is on; two people speak and the lines arrive with speaker labels.
- Speaker 0 is renamed to Ben (DM) and every line updates; F9 drops a bookmark into the journal.
- Stop: a Summary page appears in the journal. Next session, the chat is asked what was decided about the duke, and quotes the transcript.
Your key, your cost
The subscription pays for Familiar: the tools, the rules engine, every update. It never pays for a second of audio. Voices, sound effects and transcription run on keys you make at the providers of your choice and paste into the Voice and Transcription tabs; each provider bills you at its own price, and Familiar adds nothing on top.
Three things Familiar never does: meter your audio, mark up your provider, or hide the price. Where a provider reports what a clip cost, the module prints it on the cell. Where it reports nothing, Familiar shows nothing rather than guessing.
Spending is held down before anything is billed. A repeated line plays from the browser cache. A sound effect is looked up in your own playlists before one is generated, and an automatic one-shot is capped at six seconds and three generated clips a minute. Searching and casting voices cost nothing; on ElevenLabs a preview is the provider's own free sample.
One key can do two jobs. An ElevenLabs voice key also generates sound effects, and so does a fal.ai image key. That cuts both ways: an image key alone is a sound-effect key, and the sound-effects chapter says how to hold it. Transcription and dictation share a key of their own. Keys stay in your Foundry client settings, in your browser; audio is made browser-direct, and no Familiar server sits in the audio path.
A few list prices, quoted with their source so you can recompute them:
- ElevenLabs text to speech1 credit per character on the standard models; the Creator plan is $22 a month for 121,000 creditsA sixty-character line is about sixty credits, so the Creator plan covers roughly two thousand such lines a month. Source
- Cartesia Sonic text to speech$5 a month for 100,000 credits on the Pro plan, about 133 minutes of speech; the free plan holds 20,000 creditsThe free plan is about 27 minutes of speech a month; Pro is about two hours. Source
- OpenAI tts-1 and tts-1-hd$15 and $30 per million charactersFifteen thousand characters of speech, a long session, is about 23 cents on tts-1 and 45 on tts-1-hd. Source
- OpenRouter hexgrad/kokoro-82m$0.62 per million charactersThe same fifteen thousand characters are about one cent. Source
- fal.ai ElevenLabs Sound Effects v2$0.002 per second of generated audioA three-second creak is about six tenths of a cent; a twenty-second loop is four cents. Source
- Deepgram Nova-3 streaming$0.0048 per minute at the current pay-as-you-go rate; the regular rate is $0.0077; a new account starts with $200 of creditA four-hour session is about $1.15 at the current rate and $1.85 at the regular one. Source
- Gladia real-time transcription$0.75 per hour on the Starter plan; a new account starts with €50 of creditA four-hour session is $3 on Starter; Gladia advertises its own starting credit as about sixty hours of real time. Source
- AssemblyAI streaming transcription$0.15 per hour on Universal-Streaming, $0.45 on Universal-3 Pro; a new account starts with $50 of creditA four-hour session is 60 cents, or $1.80 on the Pro model; the clock runs while the connection is open, idle time included. Source
Prices checked on 28 August 2026 on the providers' public pricing pages.
Providers and settings
Speech runs on the provider you bring, and what a line can do follows that choice: Familiar sends only the knobs a provider understands and logs the ones it skips. Transcription has its own three providers, with their own key.
Speech, by provider:
| Provider | Default model | Emotion hint | Pause before a line | Speed | Volume | Language per line | Voice catalogue and preview |
|---|---|---|---|---|---|---|---|
| ElevenLabs | eleven_flash_v2_5 | On eleven_v3 only, which also sings (experimental, and it varies by voice) | Yes | No | No | Yes | Full catalogue with labels; a free sample on every stock voice |
| Cartesia | sonic-3.5 | Yes, English, from a fixed list | Yes | 0.6 to 1.5 | 0.5 to 2 | Yes, with locale | Full catalogue with labels; a sample on about a quarter of the voices |
| OpenAI TTS | gpt-4o-mini-tts | Yes, as an instruction | No, and the reply says so | 0.25 to 4 | No | No | Thirteen fixed voices; a preview is generated |
| OpenRouter TTS | hexgrad/kokoro-82m | On OpenAI models only | No | Yes | No | No | Type a voice name; a preview is generated |
| NanoGPT TTS | Kokoro-82m | Yes, as an instruction | No | Yes | No | No | Type a voice name; billed from your NanoGPT balance; a preview is generated |
| Custom OpenAI-compatible | the model you name | Yes, as an instruction | No | Yes | No | No | Type a voice name; a preview is generated |
Transcription, by provider:
| Provider | Model | Speaker labels | Languages | To start |
|---|---|---|---|---|
| Gladia | Solaria-1, the default | No | Detected automatically, switches mid-sentence | €50 of credit |
| Deepgram | Nova-3 | Yes, numbered | Many | $200 of credit |
| AssemblyAI | Universal-3 Pro streaming | Yes, lettered | English, Spanish, German, French, Portuguese, Italian | $50 of credit |
- Voice starts switched on, per browser, and stays silent and unbilled until that browser holds a key. The Voice tab holds the switch, the narrator and character slots with their providers, keys and models, Find a voice, Voice Assignments, and the Speech section: Speak aloud, Playback, Chat log, Table chat, Cast voices, Speech bubbles, Voice volume, Progressive, and the two dip rows, Music under voice and Ambient sounds.
- Player Broadcast, Ambience and Sound effects sit in the same tab. Broadcast is one select: Off, Manual or Automatic. Mood switching, under Ambience, and Sound effects each take Off, On request or Automatic. All three are world settings and all three default to Automatic, so out of the box your players hear what Familiar makes and a generated effect bills the key you pasted; each select's Off and On request rows hold it. The two dip rows are each listener's own and also appear in Foundry's Module Settings, where a player sets theirs.
- Transcription has its own tab: Enable Transcription, Enable Dictation with its send mode and pause, the provider, key and model the two share, and After the Session: searchable by the AI (on), a recap page (off), players read along (off).
- Keys are stored in your Foundry client settings, in your browser, and never pass through a Familiar server.
- Keyboard: Stop All Familiar Audio is unbound until you bind it, F9 bookmarks the transcript, Ctrl+F searches the panel, Ctrl+Shift+D is push-to-talk dictation. The Stop button also answers your keyboard's media keys.
- Over MCP the same tools run, and the clip plays in the Foundry tab, never in the client.
Works in Foundry v13 and v14, verified with v14.367. Model choice for the chat side lives in the model guide; the tool catalogue lives on the capabilities page.
Questions before you switch sound on
Does Familiar make music?
No, and that is a decision rather than a gap. Music comes from your own playlists: a pack from the Foundry package browser, your own files through Bulk Import, or an adventure module's playlists. Familiar finds the mood by name, crossfades, and binds it to the scene; the mechanism is above. Sound effects are the one thing it generates: short one-shots, and loops when you ask, never a track of music.
Does it copy my voice, or anyone else's?
No. Familiar has no voice copying and no speech-to-speech mode, on purpose: the legal ground under a copied voice is thin, and a table does not need it. Voices come from a provider's catalogue, or from a voice id you paste for a voice you already own at that provider.
Do my players need a key, or a module?
Neither. Voice clips and sound effects reach them through Foundry's own audio, under their own interface volume, and the spoken line shows on the token. They install nothing and pay nothing; what they see and hear is its own chapter.
What appears in my world?
Two playlists once a clip is filed, Familiar Ambience and Familiar Sound Effects. Two folders under the world's data, familiar-tts for broadcast clips and familiar-sfx for effects, one file per unique clip. A [Familiar] Transcripts journal folder once you transcribe. And a voice flag on each actor you cast. Nothing leaves the world, and nothing is uploaded anywhere but your own Foundry server.
Does it delete old clips?
No. Foundry gives a module no way to delete a file, so Familiar keeps no timer that would pretend otherwise, for familiar-tts and familiar-sfx alike. Clips are named by their content, so the same line in the same voice reuses its file; both folders only grow, so prune them by hand.
Why does my player hear nothing?
Browsers block audio until the page has been clicked or a key pressed, and Foundry reports that lock. One click or keypress in the Foundry window lifts it, and the next clip plays. Still silent? Check that Player Broadcast is not Off, and that the player's interface volume in Foundry is up.
Does it work on Forge, or over plain http?
Voices, ambience and effects work anywhere Foundry runs. Two things need a secure page, https or Foundry opened on localhost on the machine that runs it: the microphone for transcription and dictation, and the browser cache that makes a repeated line free. On plain http the line is generated again instead.
Does it work from Claude Desktop, or another MCP client?
Yes, with the same tools, and the audio plays in the Foundry tab rather than in the client. Three things to know. Claude Desktop's curated tool set leaves the voice bundle out, so add FAMILIAR_TOOL_BUNDLES=voice-generation to its config. Speak aloud, Ambience and Sound effects are read from get-world-info, so an MCP client follows the same three switches the built-in chat does. And with Playback on Wait for my go, an MCP client's result comes back with the narration held and a holdId, which play-held-voice and discard-held-voice act on; the chat card appears at your Play, not before.
Which voice provider should I start with?
ElevenLabs has the widest catalogue and a free sample on every stock voice. Cartesia and OpenAI both take an emotion hint on any voice, Cartesia in English, OpenAI as an instruction; OpenAI bills per character while Cartesia runs on monthly credits with a free tier. OpenRouter reaches Kokoro on a key you may already use for chat, and NanoGPT bills speech against the same balance as its chat models. The table above says what each one can do, and the dated prices say what a session costs.
Try it at your next session
Two weeks free, then $6 per month or $48 per year. Your voice, sound and transcription providers bill separately, at their own prices. Everything else Familiar does at the table is on the capabilities page.