AnkiWeb
- Rating
- 0 (π 0 Β· π 0)
- Updated
- 2026-07-01
- Anki versions
- 25.09.4~
- Description language
- en
AnkiWeb addon 1279936795
Synthesizes Japanese audio at review time via a local VOICEVOX server, with central preset routing, text cleanup rules, caching, and optional audio export.
Open on AnkiWeb GitHub Ask about alternatives
active
| Min Anki | Max Anki | Updated |
|---|---|---|
| 2.1.66 | 25.09.4+ | 2026-07-01 |
Loadingβ¦
Fetches examples from massif.la, connects to a local TTS server to generate audio, and updates hardcoded deck fields with text and audio via hotkeys, mainly for personal use.
Generates contextual written explanations and spoken audio for target words in Anki sentence cards using OpenAI models and selectable TTS engines.
Generates sentence audio via local AI TTS and translates selected Browse notes into target fields, requiring configured endpoints, mappings, and a running TTS server.
Generates high-quality Japanese audio for Anki cards via VOICEVOX, requiring the engine running and a right-click card browser workflow.
Pronounces selected text or a fallback field via Alt+C using Azure/Edge neural voices, automatic language detection, and local audio caching.
Flips bilingual Anki decks into the opposite language direction and generates offline pronunciation audio using local models, without modifying original decks.
An Anki add-on that synthesizes audio at review time through local AI TTS engines β no pre-generated media files, no cloud. Built for Japanese decks first; language-agnostic underneath.
β οΈ Current state: one provider, Japanese-focused. Only VOICEVOX is wired up so far. The architecture is provider-agnostic (Piper / Style-Bert-VITS2 / generic HTTP are designed to slot in as new files under
providers/), but nothing else is implemented yet.Heavily vibe-coded. Blame Claude if something's not awesome.
PRs welcome β fork ahead. I built this for myself and have no plans to actively maintain it. If you add a provider, fix a bug, or polish the UI, send a PR and I'll happily merge. Just don't count on me shipping updates on any schedule.
The two big existing TTS add-ons β AwesomeTTS and HyperTTS β both work, but:
Local AI engines β VOICEVOX, Piper, Style-Bert-VITS2 β sound dramatically better for Japanese, run on your own machine, and cost nothing. This addon is built around that constraint: local providers only, presets are first-class, routing is central.
Card templates use Anki's built-in TTS tag with a custom voice name:
{{tts ja_JP voices=LocalTTS:Expression}}
At review time:
Anki reviewer
β TTSProcessPlayer
Routing table deck β notetype β language β default
β
Cleanup pipeline HTML, ruby, brackets, cloze, CJK spaces
β
Regex rules global, with optional per-preset override
β
Cache lookup sha256(provider+options β processed text)
β miss
Provider VOICEVOX (Piper, Style-Bert-VITS2 next)
β
Cache write + playback Opus via ffmpeg, WAV fallback
Templates only declare that a field wants TTS. Which preset speaks is decided by a central routing table β switching voices for one deck (or your whole collection) is a single dropdown change, not a template rewrite.
docker run --rm -p 50021:50021 voicevox/voicevox_engine:cpu-latest
For other install options (native Mac app, Windows installer, GPU image, etc.) see VOICEVOX's official site at voicevox.hiroshiba.jp and the engine repo at github.com/VOICEVOX/voicevox_engine. Whatever you run, point the addon at it via Settings β Providers β endpoint (default http://localhost:50021).Easiest: from AnkiWeb via Tools β Add-ons β Get Add-ons in Anki.
Or grab the .ankiaddon from the latest GitHub release and double-click to install.
uv sync --extra dev
uv run python scripts/dev_link.py
Then restart Anki. The script symlinks local_tts/ into Anki's addons21/ folder. Edit files in the repo, restart Anki to reload.
uv run python scripts/dev_link.py --unlink removes it.
.ankiaddonuv run python scripts/build_addon.py
Produces dist/local_tts.ankiaddon β double-click to install in Anki.
{{tts ja_JP voices=LocalTTS:Expression}}
Expression is a placeholder for whichever field on your note type holds the text to be spoken β it must be the exact name of an existing field. If your note type has a field called Sentence instead, write β¦voices=LocalTTS:Sentence. The field should contain Japanese text (or whatever language matches the preset routed for it). If the field is empty for a given card, nothing plays β no error.A quick switcher is available at Tools β Local TTS Β· Routes for changing the default / per-language / per-deck preset without opening Settings.
Everything lives under Tools β Local TTS settingsβ¦. Save applies the changes on the next playback β no restart.
PATH and common install locations (Homebrew, /usr/local/bin, /usr/bin). Set explicitly if you have multiple installs. Without ffmpeg the cache falls back to WAV (~12Γ larger but still works).<addon-folder>/user_files/cache/. LRU-evicted by file access time when full.Provider-level settings, shared across every preset that uses that provider. Edit once when you move the server; presets don't need editing.
http://localhost:50021 by default. Change if you run VOICEVOX on a different host/port (e.g. http://macmini.local:50021 for a LAN server).http://localhost:50121 by default. Nemo is a sibling engine from the VOICEVOX team with character-less, neutral-sounding voices aimed at business / educational TTS. Same HTTP API as regular VOICEVOX, ships as its own binary on a separate port. Boot both engines if you want to mix regular character voices with Nemo's neutral voices in the same collection β routing handles the dispatch.A preset is a voice configuration β provider + voice ID + per-voice options (speed, pitch, β¦). Add, edit, duplicate, or delete presets here. Each preset has:
Global text and prosody settings. Apply to every preset that doesn't explicitly override them.
Fixed-order pipeline applied before regex rules. Configurable:
<ruby>ζ¬<rt>γ»γ</rt></ruby>) β keep the base character (ζ¬) or the reading (γ»γ).ζ₯γ
[γ²γ³]) β keep the base, the reading, or both.[], ().ζ₯ζ¬γγ« β ζ₯ζ¬γ«. Only collapses where both neighbours are Japanese.Always-on cleanup (not configurable): HTML strip, Anki cloze braces ({{c1::X}} β X), whitespace normalize.
1990 β εδΉηΎδΉε, 7ζ β δΈζ, 100ε β ηΎε). Without this, VOICEVOX reads bare digits one at a time ("nana tsuki" instead of "shichi-gatsu"). Off β digits are sent through verbatim.Ordered list of pattern β replacement substitutions applied after cleanup. Use them for vocabulary fixes the engine gets wrong:
20ζ₯ β γ―γ€γθθ² γ£γ¦γγ β γγγ£γ¦γγPer row: On (enable), Pattern (Python regex), Replacement (literal or backreferences). The Validate button compiles every pattern and reports failures inline. Validation also runs at config load β broken patterns produce a one-time popup and are skipped at runtime, never crashing playback.
Insert the marker character in card text where you want a short pause instead of VOICEVOX's default ~0.15s comma pause. The engine still pronounces neighbouring digits separately (δΈγ»εε β "san, yon-bai", not "sanjuuyon-bai") and prosody flows continuously across the join.
γ». Leave empty to disable entirely.0.03 s. Set 0 for no audible gap.γ sitting between two digits is rewritten to the marker before synthesis, so you don't have to type the marker in cards. Covers half-width (2023γ2024), full-width (οΌγοΌ), and CJK digits (δΈγε, εδΈγε, δΈεγεε). Commas in non-digit contexts (δΈζγεζ, ζ₯ζ¬γθ±θͺ) are untouched. Off by default.Global baseline for the per-voice numeric parameters. Each preset's editor shows an "Use global (X)" checkbox per parameter β checked = inherit, unchecked = preset pins its own value.
Changing a global value here re-synthesizes any preset that inherits it on next play. Presets that override the value are unaffected.
Decides which preset plays when a TTS tag fires. Resolution order, first match wins:
deck β preset. Most specific.notetype β preset.ja β preset, en β preset, β¦ Use the ja_JP / en_US style in your card template's TTS tag; the language root (ja) is what matches here.Tools β Local TTS Β· Routes gives a one-click submenu to switch the default / per-language / per-deck preset without opening Settings.
Off by default. The runtime cache is deliberately private β but AnkiMobile / AnkiDroid can't run a local engine at all, so the only way to get audio there is to persist a real media file into a note field and let Anki sync it. This is the sanctioned opt-in exception to the "never touches collection.media" rule.
source field β audio field in the table below. Both columns are dropdowns populated from that note type's actual field names β there's no free-text entry, so a misspelled field name is impossible to save in the first place. If a previously-saved mapping references a field that was since renamed or removed, that row shows "β name (missing)" so you can spot and fix it immediately. Scoping by note type also means the same generic field name (sentence_jp1) on two unrelated note types can never cross-wire.sentence_jp2's audio from ever landing in sentence_jp1_audio, even with several example-sentence fields sharing {{tts}} tags: if two mapped fields happen to hold identical text on the same card, the match is ambiguous and nothing is written (safe by default, never guesses). A brief on-screen message confirms each auto-save as it happens, and a separate warning appears (once per note type per session) if the configured mapping references a field the note type doesn't actually have.Ctrl+Shift+O. Qt automatically swaps CtrlβCmd on macOS, so this is Cmd+Shift+O there and literal Ctrl+Shift+O on Windows/Linux. Bound as an application-wide shortcut so it fires even while the review webview has keyboard focus. Leave the field empty to disable it.ffmpeg) rather than reused as-is from the cache. This matters because the live-playback cache's Opus files use an Ogg container, and iOS cannot decode Opus-in-Ogg at all β only Opus wrapped in Apple's own .caf container, which this addon doesn't produce β so a synced Ogg-Opus file silently fails to play on AnkiMobile. MP3 is the one format guaranteed to decode everywhere: desktop, AnkiWeb's browser <audio> element, AnkiDroid, and AnkiMobile. If ffmpeg isn't available, the source file (WAV, which already plays everywhere) is stored as-is instead.<addon-folder>/user_files/log/local_tts.log β including why nothing was saved (feature disabled, no field mapping for this note type, no unique text match, misconfigured field, field already filled, etc.). Check there first if a field isn't getting filled. Remember that Anki only reloads addon code at startup, so after upgrading the addon, fully restart Anki before testing.{{tts}} tag in a conditional so desktop doesn't play both the live-synthesized and the stored audio back to back:{{^sentence_jp1_audio}}{{tts ja_JP voices=LocalTTS:sentence_jp1}}{{/sentence_jp1_audio}}{{sentence_jp1_audio}}
This plays the live tag only while sentence_jp1_audio is empty; once it's filled, the stored [sound:...] plays instead β the same audio on desktop and mobile.<addon-folder>/user_files/cache/. Never touches collection.media.sha256(preset.fingerprint() β processed_text). The fingerprint covers only provider + options. Endpoint, cleanup flags, and regex rules are not in it β moving servers or tweaking rules doesn't invalidate audio for unaffected text.ffmpeg is on PATH (or set explicitly in Settings β General), WAV otherwise.atime, capped at the configured MB ceiling.If a provider is unreachable, you'll see a transient tooltip (e.g. Local TTS: VOICEVOX is not reachable at http://localhost:50021 β start the engine or update the endpoint.) β once per session per error type, not every card. The log at <addon-folder>/user_files/log/local_tts.log has the full picture.
local_tts/
βββ __init__.py entry point
βββ addon.py composition root, menu wiring, apply_config
βββ config.py persisted settings + load/save + addon_dir resolution
βββ presets.py Preset / RegexRule / CleanupOptions + fingerprint
βββ routing.py deck > notetype > language > default
βββ cache.py disk cache + Opus transcode + LRU eviction
βββ player.py TTSProcessPlayer subclass, _on_done queue inject
βββ export.py opt-in: match played text to a note field, save to collection.media
βββ providers/ abstract Provider + adapters
β βββ base.py
β βββ voicevox.py
βββ text/
β βββ cleanup.py pure functions, unit-testable
β βββ regex_rules.py validate_pattern + apply
βββ gui/
βββ settings.py tabbed dialog
βββ preset_editor.py modal editor + live voice picker
Provider adapters are stateless: they receive (text, preset, provider_settings) and return WAV bytes. Nothing outside providers/ may import a concrete provider.
uv sync --extra dev
uv run pytest # 130 tests, fully pure
uv run python scripts/dev_link.py # symlink into Anki
uv run python scripts/build_addon.py # build dist/local_tts.ankiaddon
Tests cover the pure modules β cleanup pipeline, preset fingerprint stability, routing precedence, cache roundtrip + Opus path, config roundtrip, regex validation. GUI and Anki-integrated code is smoke-tested manually.
pkill -x Anki && open -a Anki if you're impatient on macOS).av_player.insert_file in _on_done β Anki's default TTSProcessPlayer does not auto-enqueue.MIT β use it, modify it, share it, sell it, do whatever. No warranty, no liability. See LICENSE.
Code design was informed by AwesomeTTS (GPLv3), but no code from AwesomeTTS is included or redistributed β only design ideas (the TTSProcessPlayer hook pattern, sanitization approach).