Local TTS for Japanese (VOICEVOX)

AnkiWeb addon 1279936795

Synthesizes Japanese audio at review time via a local VOICEVOX server, with central preset routing, text cleanup rules, caching, and optional audio export.
AI-generated summary; may contain mistakes.

card-creationexportsync-integrationformattingshortcutsmobile-compatiblebeta

Open on AnkiWeb GitHub Ask about alternatives

AnkiWeb

Rating
0 (πŸ‘ 0 Β· πŸ‘Ž 0)
Updated
2026-07-01
Anki versions
25.09.4~
Description language
en

Maintenance

active

  • Last update or commit was 92 days before the snapshot (2026-07-01).
  • The repository has 10 test files.
  • The repository has 2 GitHub Actions workflows.

Will it work on my Anki?

Version branches
Min AnkiMax AnkiUpdated
2.1.6625.09.4+2026-07-01
History across monthly snapshots

Loading…

Similar addons

AI Language Explainer

Generates contextual written explanations and spoken audio for target words in Anki sentence cards using OpenAI models and selectable TTS engines.

card-creationaisync-integration

Rating 10 πŸ‘ 10 Β· πŸ‘Ž 0 ⭐ 9 Anki 2.1.66~ Updated 2025-10-07

README

Local TTS for Anki

An Anki add-on that synthesizes audio at review time through local AI TTS engines β€” no pre-generated media files, no cloud. Built for Japanese decks first; language-agnostic underneath.

⚠️ Current state: one provider, Japanese-focused. Only VOICEVOX is wired up so far. The architecture is provider-agnostic (Piper / Style-Bert-VITS2 / generic HTTP are designed to slot in as new files under providers/), but nothing else is implemented yet.

Heavily vibe-coded. Blame Claude if something's not awesome.

PRs welcome β€” fork ahead. I built this for myself and have no plans to actively maintain it. If you add a provider, fix a bug, or polish the UI, send a PR and I'll happily merge. Just don't count on me shipping updates on any schedule.

Why this exists

The two big existing TTS add-ons β€” AwesomeTTS and HyperTTS β€” both work, but:

  • Configuration is painful. Voice/preset is encoded in the card template, so changing engine or voice means re-editing every note type and re-saving. Preset management is rigid; switching providers across a deck is a chore.
  • The good voices are behind paywalls. ElevenLabs, Azure, OpenAI TTS, Google Cloud TTS, etc. all charge per character. For daily Anki reviews this adds up fast, and cloud engines also handle Japanese poorly compared to dedicated Japanese models.

Local AI engines β€” VOICEVOX, Piper, Style-Bert-VITS2 β€” sound dramatically better for Japanese, run on your own machine, and cost nothing. This addon is built around that constraint: local providers only, presets are first-class, routing is central.


How it works

Card templates use Anki's built-in TTS tag with a custom voice name:

{{tts ja_JP voices=LocalTTS:Expression}}

At review time:

Anki reviewer
   ↓  TTSProcessPlayer
Routing table             deck β†’ notetype β†’ language β†’ default
   ↓
Cleanup pipeline          HTML, ruby, brackets, cloze, CJK spaces
   ↓
Regex rules               global, with optional per-preset override
   ↓
Cache lookup              sha256(provider+options β€– processed text)
   ↓  miss
Provider                  VOICEVOX (Piper, Style-Bert-VITS2 next)
   ↓
Cache write + playback    Opus via ffmpeg, WAV fallback

Templates only declare that a field wants TTS. Which preset speaks is decided by a central routing table β€” switching voices for one deck (or your whole collection) is a single dropdown change, not a template rewrite.


Requirements

  • Anki 24.x or newer (Qt 6 / PyQt 6, Python 3.13 bundled).
  • macOS, Windows, or Linux desktop. AnkiMobile / AnkiWeb can't run local engines, so this is desktop-only.
  • A local TTS engine. VOICEVOX is the first-class provider; others come next. Quickest setup is Docker:
    docker run --rm -p 50021:50021 voicevox/voicevox_engine:cpu-latest
    
    For other install options (native Mac app, Windows installer, GPU image, etc.) see VOICEVOX's official site at voicevox.hiroshiba.jp and the engine repo at github.com/VOICEVOX/voicevox_engine. Whatever you run, point the addon at it via Settings β†’ Providers β†’ endpoint (default http://localhost:50021).
  • ffmpeg (optional but recommended). Used to transcode the cache to Opus (~12Γ— smaller than WAV). Without it the addon falls back to WAV.

Install

Easiest: from AnkiWeb via Tools β†’ Add-ons β†’ Get Add-ons in Anki.

Or grab the .ankiaddon from the latest GitHub release and double-click to install.

Run from this repo (developer)

Symlink (live edits)

uv sync --extra dev
uv run python scripts/dev_link.py

Then restart Anki. The script symlinks local_tts/ into Anki's addons21/ folder. Edit files in the repo, restart Anki to reload.

uv run python scripts/dev_link.py --unlink removes it.

Build a .ankiaddon

uv run python scripts/build_addon.py

Produces dist/local_tts.ankiaddon β€” double-click to install in Anki.


Quickstart

  1. Start VOICEVOX (or your engine of choice).
  2. Install the addon (above).
  3. Tools β†’ Local TTS settings…
    • Providers tab: confirm or change the VOICEVOX endpoint.
    • Presets tab: the bundled "Japanese VOICEVOX ζ˜₯ζ—₯ιƒ¨γ€γ‚€γŽ Β· γƒŽγƒΌγƒžγƒ«" preset is ready. Click New to add more; the editor's "Pick voice from server…" button (only on new presets) fetches the live speaker list from VOICEVOX with a built-in test-play. Each preset may optionally override the global cleanup / regex rules.
    • Rules tab: global text cleanup, regex substitutions, split marker for tight pauses, and voice defaults inherited by every preset. See the Settings guide for details.
    • Routing tab: which preset plays for which deck / note type / language.
  4. Card template β€” add to any field you want spoken. Edit your note type's card template (Browse β†’ Cards…) and insert, for example:
    {{tts ja_JP voices=LocalTTS:Expression}}
    
    Expression is a placeholder for whichever field on your note type holds the text to be spoken β€” it must be the exact name of an existing field. If your note type has a field called Sentence instead, write …voices=LocalTTS:Sentence. The field should contain Japanese text (or whatever language matches the preset routed for it). If the field is empty for a given card, nothing plays β€” no error.
  5. Review a card. Cache miss β†’ ~1–2s synth + play. Cache hit β†’ instant.

A quick switcher is available at Tools β†’ Local TTS Β· Routes for changing the default / per-language / per-deck preset without opening Settings.


Settings guide

Everything lives under Tools β†’ Local TTS settings…. Save applies the changes on the next playback β€” no restart.

General

  • Enable Local TTS β€” master switch. Off leaves the addon installed but inert.
  • Default preset β€” used when no per-deck / per-notetype / per-language rule matches.
  • ffmpeg path β€” leave blank to auto-detect on PATH and common install locations (Homebrew, /usr/local/bin, /usr/bin). Set explicitly if you have multiple installs. Without ffmpeg the cache falls back to WAV (~12Γ— larger but still works).
  • Cache size limit β€” total disk budget for synthesized audio under <addon-folder>/user_files/cache/. LRU-evicted by file access time when full.
  • Clear cache now β€” delete every cached file. Next playback will re-synthesize on demand.

Providers

Provider-level settings, shared across every preset that uses that provider. Edit once when you move the server; presets don't need editing.

  • VOICEVOX β†’ endpoint β€” http://localhost:50021 by default. Change if you run VOICEVOX on a different host/port (e.g. http://macmini.local:50021 for a LAN server).
  • VOICEVOX Nemo β†’ endpoint β€” http://localhost:50121 by default. Nemo is a sibling engine from the VOICEVOX team with character-less, neutral-sounding voices aimed at business / educational TTS. Same HTTP API as regular VOICEVOX, ships as its own binary on a separate port. Boot both engines if you want to mix regular character voices with Nemo's neutral voices in the same collection β€” routing handles the dispatch.

Presets

A preset is a voice configuration β€” provider + voice ID + per-voice options (speed, pitch, …). Add, edit, duplicate, or delete presets here. Each preset has:

  • Name β€” what shows up in the routing table.
  • Provider β€” which engine to use. Currently VOICEVOX or VOICEVOX Nemo.
  • Pick voice from server… (new-preset only) β€” live-queries VOICEVOX for the speaker list with a built-in test-play button.
  • Speaker / speed / pitch / intonation / volume β€” voice parameters. By default these inherit from Voice defaults under the Rules tab; tick the "Use global" checkbox off if you want this preset to pin its own value.
  • Override global cleanup / regex rules β€” opt-in checkboxes. Off (default) means the preset inherits the global Rules. On replaces the global list with the preset's own block.

Rules

Global text and prosody settings. Apply to every preset that doesn't explicitly override them.

Cleanup

Fixed-order pipeline applied before regex rules. Configurable:

  • Ruby tags (<ruby>本<rt>ほん</rt></ruby>) β†’ keep the base character (本) or the reading (ほん).
  • Bracket readings (ζ—₯γ€…[ひび]) β†’ keep the base, the reading, or both.
  • Bracket pairs β€” which bracket characters get the bracket-reading treatment. Defaults: [], ().
  • Collapse spaces between Japanese characters β€” removes furigana-add-on artifacts like ζ—₯ζœ¬γ€€γ« β†’ ζ—₯本に. Only collapses where both neighbours are Japanese.

Always-on cleanup (not configurable): HTML strip, Anki cloze braces ({{c1::X}} β†’ X), whitespace normalize.

Numbers

  • Read numbers as words (on by default) β€” rewrites every run of ASCII or full-width digits to its classical kanji form before synthesis (1990 β†’ 千九百九十, 7月 β†’ δΈƒζœˆ, 100円 β†’ 百円). Without this, VOICEVOX reads bare digits one at a time ("nana tsuki" instead of "shichi-gatsu"). Off β†’ digits are sent through verbatim.

Regex rules

Ordered list of pattern β†’ replacement substitutions applied after cleanup. Use them for vocabulary fixes the engine gets wrong:

  • 20ζ—₯ β†’ は぀か
  • θƒŒθ² γ£γ¦γγ‚‹ β†’ しょってくる

Per row: On (enable), Pattern (Python regex), Replacement (literal or backreferences). The Validate button compiles every pattern and reports failures inline. Validation also runs at config load β€” broken patterns produce a one-time popup and are skipped at runtime, never crashing playback.

Split marker (VOICEVOX)

Insert the marker character in card text where you want a short pause instead of VOICEVOX's default ~0.15s comma pause. The engine still pronounces neighbouring digits separately (三・四倍 β†’ "san, yon-bai", not "sanjuuyon-bai") and prosody flows continuously across the join.

  • Marker character β€” what to look for in card text. Default ・. Leave empty to disable entirely.
  • Pause length β€” gap inserted at marker positions. Default 0.03 s. Set 0 for no audible gap.
  • Auto-mark digit-、-digit pauses β€” when on, any 、 sitting between two digits is rewritten to the marker before synthesis, so you don't have to type the marker in cards. Covers half-width (2023、2024), full-width (1、2), and CJK digits (三、四, 十三、四, 三十、四十). Commas in non-digit contexts (δΈ‰ζœˆγ€ε››ζœˆ, ζ—₯ζœ¬γ€θ‹±θͺž) are untouched. Off by default.

Voice defaults

Global baseline for the per-voice numeric parameters. Each preset's editor shows an "Use global (X)" checkbox per parameter β€” checked = inherit, unchecked = preset pins its own value.

  • speed β€” 1.0 is natural pace. <1 slower, >1 faster. Recommended 0.9–1.2.
  • pitch β€” offset from the speaker's natural pitch. 0.0 = unchanged. VOICEVOX is sensitive here β€” stay within roughly Β±0.1.
  • intonation β€” how much the pitch moves while speaking. 1.0 normal, 0.0 monotone/robotic, >1 exaggerated. Recommended 0.8–1.3.
  • volume β€” output multiplier. 1.0 default, 1.5–2.0 audibly louder, above ~2.0 starts clipping.

Changing a global value here re-synthesizes any preset that inherits it on next play. Presets that override the value are unaffected.

Routing

Decides which preset plays when a TTS tag fires. Resolution order, first match wins:

  1. By deck β€” deck β†’ preset. Most specific.
  2. By note type β€” notetype β†’ preset.
  3. By language β€” ja β†’ preset, en β†’ preset, … Use the ja_JP / en_US style in your card template's TTS tag; the language root (ja) is what matches here.
  4. Default preset β€” set under the General tab.

Tools β†’ Local TTS Β· Routes gives a one-click submenu to switch the default / per-language / per-deck preset without opening Settings.

Audio export

Off by default. The runtime cache is deliberately private β€” but AnkiMobile / AnkiDroid can't run a local engine at all, so the only way to get audio there is to persist a real media file into a note field and let Anki sync it. This is the sanctioned opt-in exception to the "never touches collection.media" rule.

  • Save generated audio to note fields β€” master switch for this feature.
  • Field mapping β€” per note type. Pick a note type from the dropdown, then map its source field β†’ audio field in the table below. Both columns are dropdowns populated from that note type's actual field names β€” there's no free-text entry, so a misspelled field name is impossible to save in the first place. If a previously-saved mapping references a field that was since renamed or removed, that row shows "⚠ name (missing)" so you can spot and fix it immediately. Scoping by note type also means the same generic field name (sentence_jp1) on two unrelated note types can never cross-wire.
  • While reviewing, the just-played field's text is matched against the source fields configured for that note's note type; if exactly one matches, its audio is written to the mapped audio field β€” but only once, while that field is still empty. This is what keeps sentence_jp2's audio from ever landing in sentence_jp1_audio, even with several example-sentence fields sharing {{tts}} tags: if two mapped fields happen to hold identical text on the same card, the match is ambiguous and nothing is written (safe by default, never guesses). A brief on-screen message confirms each auto-save as it happens, and a separate warning appears (once per note type per session) if the configured mapping references a field the note type doesn't actually have.
  • "Local TTS: save audio for this card" β€” synthesizes every mapped field on the current note in one press and overwrites its audio field, regardless of whether it was already filled. Use this to backfill a card you've already reviewed, one whose sentence was never actually played (e.g. hidden behind a "show more" toggle), or to refresh audio after changing a voice. Three ways to trigger it:
    • Tools menu β€” "Local TTS: save audio for this card" (works while reviewing).
    • Keyboard shortcut β€” configurable in Settings β†’ Audio export, default Ctrl+Shift+O. Qt automatically swaps Ctrl↔Cmd on macOS, so this is Cmd+Shift+O there and literal Ctrl+Shift+O on Windows/Linux. Bound as an application-wide shortcut so it fires even while the review webview has keyboard focus. Leave the field empty to disable it.
    • πŸ”Š button in the note editor toolbar β€” works in the Browser, the Add dialog, and the Reviewer's answer-side editor, operating on whichever note is currently open (not just the reviewed card).
  • Saved files are transcoded to MP3 (64 kbps, requires ffmpeg) rather than reused as-is from the cache. This matters because the live-playback cache's Opus files use an Ogg container, and iOS cannot decode Opus-in-Ogg at all β€” only Opus wrapped in Apple's own .caf container, which this addon doesn't produce β€” so a synced Ogg-Opus file silently fails to play on AnkiMobile. MP3 is the one format guaranteed to decode everywhere: desktop, AnkiWeb's browser <audio> element, AnkiDroid, and AnkiMobile. If ffmpeg isn't available, the source file (WAV, which already plays everywhere) is stored as-is instead.
  • Troubleshooting: every branch of the auto-fill and manual-save logic logs at DEBUG/INFO to <addon-folder>/user_files/log/local_tts.log β€” including why nothing was saved (feature disabled, no field mapping for this note type, no unique text match, misconfigured field, field already filled, etc.). Check there first if a field isn't getting filled. Remember that Anki only reloads addon code at startup, so after upgrading the addon, fully restart Anki before testing.
  • Avoiding double playback on desktop: once a field has a stored sound tag, wrap the live {{tts}} tag in a conditional so desktop doesn't play both the live-synthesized and the stored audio back to back:
    {{^sentence_jp1_audio}}{{tts ja_JP voices=LocalTTS:sentence_jp1}}{{/sentence_jp1_audio}}{{sentence_jp1_audio}}
    
    This plays the live tag only while sentence_jp1_audio is empty; once it's filled, the stored [sound:...] plays instead β€” the same audio on desktop and mobile.

Cache

  • Location: <addon-folder>/user_files/cache/. Never touches collection.media.
  • Key: sha256(preset.fingerprint() β€– processed_text). The fingerprint covers only provider + options. Endpoint, cleanup flags, and regex rules are not in it β€” moving servers or tweaking rules doesn't invalidate audio for unaffected text.
  • Format: Opus if ffmpeg is on PATH (or set explicitly in Settings β†’ General), WAV otherwise.
  • Eviction: LRU by file atime, capped at the configured MB ceiling.
  • Clear: Settings β†’ General β†’ "Clear cache now".

Failure handling

If a provider is unreachable, you'll see a transient tooltip (e.g. Local TTS: VOICEVOX is not reachable at http://localhost:50021 β€” start the engine or update the endpoint.) β€” once per session per error type, not every card. The log at <addon-folder>/user_files/log/local_tts.log has the full picture.


Architecture, briefly

local_tts/
β”œβ”€β”€ __init__.py            entry point
β”œβ”€β”€ addon.py               composition root, menu wiring, apply_config
β”œβ”€β”€ config.py              persisted settings + load/save + addon_dir resolution
β”œβ”€β”€ presets.py             Preset / RegexRule / CleanupOptions + fingerprint
β”œβ”€β”€ routing.py             deck > notetype > language > default
β”œβ”€β”€ cache.py               disk cache + Opus transcode + LRU eviction
β”œβ”€β”€ player.py              TTSProcessPlayer subclass, _on_done queue inject
β”œβ”€β”€ export.py              opt-in: match played text to a note field, save to collection.media
β”œβ”€β”€ providers/             abstract Provider + adapters
β”‚   β”œβ”€β”€ base.py
β”‚   └── voicevox.py
β”œβ”€β”€ text/
β”‚   β”œβ”€β”€ cleanup.py         pure functions, unit-testable
β”‚   └── regex_rules.py     validate_pattern + apply
└── gui/
    β”œβ”€β”€ settings.py        tabbed dialog
    └── preset_editor.py   modal editor + live voice picker

Provider adapters are stateless: they receive (text, preset, provider_settings) and return WAV bytes. Nothing outside providers/ may import a concrete provider.


Development

uv sync --extra dev
uv run pytest                                 # 130 tests, fully pure
uv run python scripts/dev_link.py             # symlink into Anki
uv run python scripts/build_addon.py          # build dist/local_tts.ankiaddon

Tests cover the pure modules β€” cleanup pipeline, preset fingerprint stability, routing precedence, cache roundtrip + Opus path, config roundtrip, regex validation. GUI and Anki-integrated code is smoke-tested manually.


Limitations

  • Live synthesis is desktop only. AnkiMobile / AnkiWeb can't run local engines. Enable Audio export and a card's audio will play on mobile too, but only for fields you've saved that way β€” mobile never synthesizes anything itself.
  • One config per Anki profile. Multi-profile users get one independent install each, which is fine.
  • Anki reloads addons only at startup. After edits, restart Anki (pkill -x Anki && open -a Anki if you're impatient on macOS).
  • The replay button works via an explicit av_player.insert_file in _on_done β€” Anki's default TTSProcessPlayer does not auto-enqueue.

License

MIT β€” use it, modify it, share it, sell it, do whatever. No warranty, no liability. See LICENSE.

Code design was informed by AwesomeTTS (GPLv3), but no code from AwesomeTTS is included or redistributed β€” only design ideas (the TTSProcessPlayer hook pattern, sanitization approach).