Export Known Words

AnkiWeb addon 596526198

Exports known words from an Anki collection as text, clipboard, or JSON, selecting deck, field, and mature, young, learning, or relearning card states.
AI-generated summary; may contain mistakes.

export

Open on AnkiWeb GitHub Ask about alternatives

AnkiWeb

Rating
0 (👍 0 · 👎 0)
Updated
2026-09-21
Anki versions
26.09~
Description language
en

Maintenance

active

  • Last update or commit was 7 days before the snapshot (2026-09-24).
  • The repository has 1 GitHub Actions workflows.

Will it work on my Anki?

Version branches
Min AnkiMax AnkiUpdated
2.1.5026.09+2026-09-21
History across monthly snapshots

Loading…

Similar addons

Deck Words Exporter

Exports vocabulary from any decks to clean .txt or .json for LLM analysis, with filters, metadata, deduplication, field selection, sorting, and configurable output.

aiexportfsrsstatisticsdeck-management

Rating 0 👍 0 · 👎 0 Anki 2.1.1~ Updated 2026-05-24
Hashi

Exports Anki review statistics and surface-word lexical snapshots from mature cards, configurable by deck and note field, with optional HTTP-triggered exports.

exportsync-integrationstatistics

Rating 0 👍 0 · 👎 0 Anki 2.1.1~ Updated 2025-12-28
Export Known Words to Clipboard

Exports known words and sentences from mature Anki cards to the clipboard for easy pasting into apps like Migaku, spreadsheets, or text editors.

export

Rating 8 👍 5 · 👎 3 Anki 25.02~ Updated 2026-03-09
Stats & Review Log Exporter

Exports collection metadata, deck summaries, optional per-card scheduling and full review log to one JSON file for external analysis.

exportstatistics

Rating 0 👍 0 · 👎 0 Anki 2.1.1~ Updated 2026-07-02

README

Torval

A pop-up dictionary for Firefox and Chrome, for Japanese, Italian and Spanish. Hold Shift and point at a word; its meaning appears next to the cursor. Everything is on your own machine, the whole dictionary lives in the browser. It works with the network off, and it collects nothing: see PRIVACY.md.

Japanese uses JMdict, the dictionary Jisho is built from. Italian and Spanish use Wiktextract, Wiktionary's own entries machine-extracted by kaikki.org, with every inflected form Wiktionary lists indexed onto the word it belongs to, so intere finds intero, capirne finds capire, and diciéndoselo finds decir. The languages are otherwise built the same way and behave the same way: pick one from the toolbar popup and its own dictionary, deinflector, known/ignored word lists and Anki settings load, independently of whatever the others have. Japanese also gets pitch accent, from Kanjium; the other two get word stress marked inline instead, since neither has lexical pitch accent to speak of. Only one language is active at a time.


Setup

Two steps, once.

1. Build a dictionary, for whichever language you want first (all three can be built; only the language picked in the toolbar popup is actually loaded).

Japanese: downloads JMdict (about 10 MB) and the JPDB/BCCWJ frequency lists, and converts them into a form the extension can read quickly.

node tools/build-dict.mjs

Then the pitch accent data, which is separate and much smaller:

node tools/build-pitch.mjs

Italian: downloads a Wiktextract dump of Italian entries (about 75 MB) and two frequency lists, hermitdave/FrequencyWords and the Wikipedia word counts, and converts them the same way. The second arrives as .xz, which Node cannot read and which is not worth a dependency for, so the build shells out to xz. Stress marks are computed as part of this step, from each entry's pronunciation, so there is no separate build for those.

node tools/build-dict-it.mjs

Spanish: the same source and the same step, about 90 MB of Wiktextract and a frequency list. Its stress marks are not read off a pronunciation but worked out from the spelling, which in Spanish says outright where the stress is; see below.

node tools/build-dict-es.mjs

2. Load it into the browser. On Firefox, go to about:debugging → This Firefox → Load Temporary Add-on… and pick extension/manifest.json. On Chrome, run node tools/package.mjs, then go to chrome://extensions, turn on Developer mode and Load unpacked the unzipped Chrome build from dist/. The toolbar popup leads with a language picker the first time; pick one to get started, or switch later from that same popup or from Settings. Choosing a language you have not built the dictionary for yet just means that language's popup waits, the same way the very first run does for whichever language you built.

The first time it runs, the extension spends a minute or so copying the dictionary into the browser's own storage: 218,000 entries and 465,000 searchable forms for Japanese, 129,000 and 586,000 for Italian, 117,000 and 764,000 for Spanish. A percentage shows on the toolbar button while it does, and until it reaches the end, hovering a word says so rather than answering.

This happens once, not once per session. It happens again only when the dictionary format changes, which the version number in meta.json decides. Progress is written down as it goes, so if it is interrupted it carries on from where it stopped rather than starting over.

Temporary add-ons are removed when Firefox restarts, so you will need to load it again each time until it is signed. See Releasing below for what that takes; node tools/package.mjs builds the file to be signed today.


Using it

Just hovering a Japanese word gives it a quiet highlight, so a page shows at a glance what Torval can help with. Click that word, or hold Shift and point at it, to actually open the dictionary.

Wherever the cursor lands inside a word finds the whole word, not just whatever happens to start under it. Point at フェ in the middle of ネカフェ and the dictionary still shows ネカフェ, not フェ on its own: Torval tries every plausible starting point behind the cursor and keeps the longest real word that actually reaches it.

point at a word a quiet highlight, no popup yet
click a word, or Shift + point at it open the dictionary
click it again close it
Shift with text selected look up the selection
Esc, a click outside, a scroll close
other matches other words that start at the same place
A / D the line before, and the line after
1 mark the word under the cursor as one you do not know
2 mark it as one you do, popup or not
3 ignore it instead: never mention this one again
B look this word up in Anki
S, or » on the bar on a video, play the stretches with nobody speaking faster
click a sense put only that meaning on the card
built from fare the word this one is made of; press it to look that one up here
+ add the word to Anki
✓ mark as known, or unmark it
⊘ ignore, or stop ignoring

Every key in that table can be changed, under Keys in the settings: click the key you want to change and press the one you want instead. The page also says which of them you have changed, and has a button to put all of them back. Torval lives on top of somebody else's page, and B is bold in every editor there is, so a shortcut that cannot be moved is one that eventually makes a page unusable.

B looks the word up in Anki itself: it opens the card browser searching for that word, in every deck and every field, and brings Anki to the front. It acts on what you have selected if you have selected something, otherwise on whatever the cursor is over. Unlike the numbers it leaves the key alone when there is nothing to look up, since a letter is a letter and plenty of sites have their own use for it.

1, 2 and 3 are taken from the page whether or not a word is under the cursor, since YouTube reads them as "jump to 30% of the video" and a key that sometimes marks a word and sometimes throws away your place is worse than either on its own.

While Shift is held the popup follows whatever you point at, and closes if you point at something that is not a word. Let go of Shift, or click a word instead, and it stays put, so you can move across and read it. A click never takes over a link, a button, a form field, or text you were dragging to select, so nothing about ordinary browsing changes.

Tags that hold for the whole word, uk, "usually written in kana", sit beside it rather than against every definition. JMdict files them per sense, but a tag on every sense is describing the word, and 事 carrying "usually kana" ten times over says nothing ten times. Where a tag really is on only some senses it stays with them: 綺麗 is usually kana when it means "clean", not when it means "pretty", and that is worth knowing. Part of speech works the same way: 勉強 is a transitive する verb for one sense and intransitive for another and just a plain noun for a third, so only "noun", what every sense actually agrees on, sits beside the word; each sense's own line carries whatever it adds beyond that.

The word a word is made of

Some words are another word with something stuck on it, and are still filed under a meaning of their own. Italian farci is fare plus ci; Wiktionary records both that and the regional sense "to simulate; to act; to pretend", and the popup used to show only the second. That is a true sentence about farci and almost never the one you need: what unlocks the line is fare, and it was nowhere on the card.

So a word like that now says built from fare, on its own line under the word and above its definitions, and the verb is pressable: it looks fare up and shows it in the same popup, in place, the way turning one page of a paper dictionary would. Which word to name is decided when the dictionary is built, from the link Wiktionary itself puts on the "compound of" sense, so it is data rather than a guess at the prose; a word whose lemma is not itself in the dictionary carries no link at all, so pressing one never lands on nothing.

This is also what already sends capirne to capire. The difference is only that capirne has no meaning of its own, so the whole word is redirected, where farci has one and keeps its own entry with a pointer on it.

Two small numbers follow the reading. The one in brackets is the pitch accent: 0 means the pitch never drops, otherwise it is the mora it drops after; the diagram of it goes on the card rather than in the popup. The other is how common the word is: 読む is top 1k, 図書館 is top 10k.

That is a band rather than a rank on purpose. A bare number asks you to know the scale already, #7,261 means nothing unless you have a feel for what #3,000 is like, and it claims a precision the data does not have. The gap between #100 and #400 is real; the gap between #7,261 and #7,800 is noise. A round band says both of those at once and needs no legend. The exact rank is on hover for when it matters. A word with no band at all is one neither corpus ever saw, which tells you something in itself.

The rank itself is blended from two corpora that read nothing alike: JPDB, built from anime, manga and visual novels, and BCCWJ, a government-run sample of newspapers, books and the web. A word common in casual speech but rare in print, or the other way round, still comes out ordinary once both are asked, rather than looking rare just because one of the two happens not to cover it.

Point at the first character of a word. Japanese has no spaces, so the extension reads forward from wherever you are pointing and finds the longest thing that is a word, point at 日 in 日本語 and you get 日本語, not 日.

Longest is not always right, and there is one well-known trap for it: a common word followed by a single particle can spell the same characters as a real, much rarer dictionary entry. 今日 ("today") plus は spells 今日は, a dated way to write こんにちは ("hello"), and JMdict really does list it. Rather than always trusting length, Torval checks whether trimming off a trailing particle lands on a dramatically more common word, and if so shows that instead. The rare reading has not gone anywhere. It sits right there under "other matches" for the rare case that is genuinely what was meant. That list is labelled "other" rather than "shorter" for exactly this reason: what shows up there is not always shorter, just not the best guess.

In Italian the same search runs across spaces as well as within a word. Nine thousand of the dictionary's entries are more than one word long, rendere conto, a meno che, pollice verso, and looking each word up alone could never find any of them. The longest phrase that really is in the dictionary wins, and gets one unbroken mark across the space; the first word on its own stays under "other matches", for when that is what was wanted.

It undoes conjugation on the way. 食べなかった is not in any dictionary, so it is walked back to 食べる and the steps taken are shown underneath, small and grey: negative → past.

Only a real na-adjective takes な or に as part of its own grammar, 元気な and 元気に are 元気 behaving adjectivally, since 元気 is tagged as one. A plain noun is not, so ネカフェに ("to the net cafe") is the word ネカフェ plus the ordinary particle に after it, not one long word ending in に; this used to be read as the latter, for any noun at all.


Anki

Cards go straight into Anki through AnkiConnect, the Anki add-on that opens a small server on your own machine. Anki has to be open; nothing leaves your computer.

Open the settings from Torval's toolbar button and choose a deck and a note type. Both lists are read from Anki itself, so a name can never be slightly wrong. Torval then fills in the field mapping by guessing from the field names, a note type with fields called Target Word, Reading, Sentence and Definitions needs no setting up at all. Anything it guesses wrongly is one dropdown away, and anything left blank stays empty on the card.

Four things can be put on a card:

Target word the dictionary form, so 食べなかった files under 食べる. An Italian or Spanish noun comes with its article on it, il cane, l’amico, el agua: the gender is half of what there is to know about a noun, and the article is the way a speaker actually carries it around
Reading the kana
Sentence the whole sentence, with the word in bold. On a video this is the whole subtitle line, not the clause the word sits in: the line was written as one thing said, and half of it on a card is half of what was said
Definition every sense, numbered
Word audio a recording of the word, if one can be found
Pitch accent the accent diagram, drawn as an SVG
Video frame the frame on screen when you pressed +
Sentence before what was said just before, for a note type that has somewhere to put it. Off unless a field asks for it
Sentence after and what was said just after
Sentence audio the subtitle line, spoken. Audio lead-in in the settings says roughly how much sound to keep from before the line, a tenth of a second by default, so the first word is not clipped by a subtitle that appears exactly as it is said. Roughly, because a recorder swallows an unpredictable moment when it starts, measured at anything from 0.05 to 0.4 seconds, so Torval starts it early and lets it

Japanese audio comes from JapanesePod101's dictionary. It answers every request with an mp3 and a 200 even when it has nothing, handing back a fixed "audio unavailable" recording instead, so Torval hashes what comes back and discards that one, leaving the field empty rather than filling your collection with identical clips.

Italian and Spanish audio comes from Lingua Libre, a Wikimedia project where volunteers read their own language a word at a time, hosted on Wikimedia Commons under CC BY-SA. There is no endpoint to ask, so the question is answered while the dictionary is built: tools/build-audio.mjs walks the category, reads the word out of each filename, and the entry ends up carrying who recorded it and where Commons keeps it. A word with no recording therefore makes no request at all, and the field is simply left empty. About 10,400 Italian words and 17,100 Spanish ones have one, weighted towards common vocabulary: a little over half the thousand commonest Italian words, three quarters of the Spanish.

That is the whole of it, and it was worth checking rather than assuming. The Lingua Libre categories hold 12,637 Italian files and 19,128 Spanish, and the build already takes every one that names a word in the dictionary. The older Category:Italian pronunciation and Category:Spanish pronunciation trees look like a second source and are not: walking them and their subcategories turns up 368 Italian words and 113 Spanish that Lingua Libre does not already have. Forvo has the coverage and a licence that forbids passing recordings on. So a word without a recording here does not have a free human one anywhere, and the field stays empty.

Where that matters, Anki says the word itself. A {{tts it_IT:Word}} line on the card template speaks whatever is in the field using the voices already on the device, which covers the words no volunteer has read yet. Torval does not write that line, because it is a decision about somebody's own note type.

Anki downloads and stores nothing itself; Torval passes it the file.

Pitch accents come from Kanjium, which derives from the NHK accent dictionary and 大辞林, the same data Yomitan and AJT Pitch Accent use. A word's whole pattern follows from one number, where the pitch drops, and the diagram is drawn from that: a dot per mora, high or low. The hollow dot on the end is the particle that would follow, which is the only thing distinguishing 橋 (pitch drops after it) from 日本語 (it does not). It is drawn in currentColor, so it takes the colour of whatever card it lands on, night mode included.

A reading alone is not enough to place an accent, 箸, 橋 and 端 are all はし with three different accents, so where the word cannot be identified the field is left empty rather than guessed at.

The bold marks the word as the page wrote it, so a conjugated form is highlighted in full, 「<b>食べなかった</b>ので、お腹が空いた。」, while the Target Word field still says 食べる.

Duplicates are allowed. A repeated word gets a small note the moment + is pressed, before the slower work of capturing the sentence even starts, but the card is made either way. One sentence often teaches several words, and mining the same word again later is not a mistake either.

By default a card carries every sense of the word. Clicking a sense before pressing + narrows it to the meanings you actually met, 語 is both "word; term" and "language", and you rarely want both. Clicking nothing means all of them, so the ordinary case needs no clicks. Picking several senses of the same word is fine, but a card is one word: choosing a sense in a different entry lets go of whatever was chosen before, rather than quietly mixing meanings from two different words onto one card.


Putting the line before or after on one card

The two fields above are all or nothing: map them and every card gets them. Usually that is not what you want, because usually a line on its own is the right amount to put on a card and more is noise you have to read every time it comes up.

So the popup offers them one card at a time. Where there is a line either side, a row appears under the entry reading also on the card: followed by the neighbouring text itself, dimmed. Click one and it is folded into the sentence for that card. Click nothing and nothing changes, which is the ordinary case. Nothing is remembered: the next word starts clean.

You can see what you are adding before you add it, which is the point. Whether a line needs its neighbour is not a rule, これはちょっと… means nothing without the question it answers and 猫です。 means everything on its own, and the only way to tell is to look.

The word stays bold in the right place, and on a caption cut in half the bold grows to the whole word once the next line completes it: 嬉しかっ becomes 嬉しかった.


Asking about a word without asking the database

Reading a page asks about far more words than it finds. Every stretch of text from every position is deinflected every way it could have been inflected, which is around thirty questions per character: on a page of thirty thousand characters, nearly nine hundred thousand of them. Nineteen in twenty are not words at all. They are the shapes a word might have taken, and the dictionary has never heard of them.

Each of those used to be a separate read of the database. Now the background script keeps a sorted list of one number per word it knows, a plain FNV hash, about two megabytes, taken once at startup. A number not in the list belongs to a word that certainly is not there, and the question is answered in memory.

Measured on that same thirty thousand characters: 883,534 reads become 14,468, one and a half percent of what it was. Ninety-four of those find nothing, because two words shared a number; there are 32 such pairs in the whole dictionary of 465,350 forms.

The same list then saves the reading a second time, and more of it. A shape the dictionary cannot have used to be held in a map, asked about, and looked for again in the answer, three pieces of work over nothing. Dropped at the moment it is proposed, a thirty thousand character page carries 22,788 terms instead of 957,322 and reads in 3.6 seconds instead of 5.9, for exactly the same 12,953 words. The tests read the same passage both ways and require the two to agree word for word.

It can only ever be wrong in the harmless direction. A shared number costs one wasted read that finds nothing, which is exactly what used to happen every time anyway. It can never say no about a word that is really there, because the number is taken from the word itself, and the tests check that against every single word in the dictionary rather than a sample.


How much of a page is read

Reading a page means segmenting every stretch of Japanese on it, which is real work: roughly a second for a long article. An article is worth it, since the number at the top is about the whole article.

A video page is not an article. Its number comes from the transcript, which Torval fetches separately, and the page around the player is comments and menus running to tens of thousands of characters, all of it read, almost none of it visible. So where the number comes from somewhere else, only what is on screen gets read, and a screen either side of it. Scrolling reads what you scrolled to, once you stop. Nothing is lost: the only reason to read a page whose score is already known is to mark the words on it, and a mark you cannot see is not doing anything.

More of a transcript arriving is likewise a reason to work out the number again and not a reason to read the page again, which used to happen every twenty seconds for the whole length of a video.

Whether the page is in the language at all

Japanese never has to be asked: a page either has Japanese characters on it or it does not. Italian shares its alphabet with the page around it, and an ordinary English page does contain Italian words, because in, a, no and ago (a needle) are all real entries in an Italian dictionary. Finding one is no evidence of anything, and the bar used to come up on every English page in the browser.

The proportion is the evidence, not the presence. Running Italian is very nearly all Italian words, around nine in ten; English prose scores a quarter of that, from the handful of short words the two languages happen to share. So a page has to be more than half recognised, over at least a few words, before Torval says anything about it. A subtitle line is exempt: the video's transcript already settled the question, and ten words are too few to settle it again.

While a page is being read, the handle counts up. It is the background script saying how far through the text it has got, each time it stops to ask the dictionary something. Several seconds of “Reading this page…” with nothing moving is indistinguishable from nothing happening.


Spanish

Spanish was added by giving the parts that were already general a language to be general about, rather than by writing a second Italian. It shares the lookup engine, the popup, the reader, the subtitle timing, the Anki side and the word lists; what is its own is a character class, a deinflector and a stress rule. Italian's files were rearranged to make that true, and they answer exactly as they did before, which the tests check on purpose.

The deinflector is a longer table than Italian's, because Spanish uses more of its verb system in ordinary speech. Italian's passato remoto is literary and deinflect-it.js leaves it out; Spanish's pretérito is how anybody says what happened yesterday, so comiste has to reach comer. Both subjunctives are in for the same reason: quiero que hables and si hablara are everyday Spanish, not a register a reader can skip.

Pronouns stuck on the end of a verb are in too, which Italian's table does not attempt. Dármelo, hablarle, diciéndoselo are each written as one word, so a reader hovering one is hovering something no dictionary lists. Attaching a pronoun also moves a written accent onto the verb (dar → dármelo), and the rules undo both at once, since the two always happen together.

What it does not do is reach inside a word. Pienso is pensar and duermo is dormir, and no suffix rule can get from one to the other. Those come from the dictionary instead, which carries every form Wiktionary actually lists pointed at the word it belongs to, 657,000 of them for Spanish. The rules are for the forms Wiktionary never wrote down; between the two, very little is missed.

Where the stress is

Italian has to be told where its stress falls, and the build reads it off a pronunciation or guesses. Spanish says so in the spelling, every time, in three rules with no exceptions in them:

  • a written accent wins outright: canción, árbol, reír
  • otherwise a word ending in a vowel, n or s is stressed on the next-to-last syllable: casa, joven, hablas
  • otherwise on the last: hablar, ciudad, feliz

That is a definition rather than a heuristic, because the accent is written exactly when the first two rules would disagree. So the bold vowel in a Spanish popup is not a best guess the way the Italian one sometimes is.

The work that is left is counting syllables, which the spelling also settles: two vowels are one syllable when at least one is an unaccented i or u (bue-no, ciu-dad, vein-te), and two otherwise (ca-er, le-al). Two letters are not vowels although they look like them: the u of que, qui, gue and gui is written but not said, which is exactly what ü exists to say otherwise, and y is a consonant except at the end of a word, where it closes a diphthong (rey, muy).


Subtitles, and mining from video

Torval times YouTube's subtitles, in whichever language is selected, so it knows exactly when each line starts and ends, that timing is what lets a line be replayed and recorded precisely, and what lets A and D jump between lines. What you see is drawn by Torval itself, in the same look as the popup.

Nothing has to be selected first. Which track to fetch is decided from the language being read, against the list of tracks the video actually has, and never from whatever the player happens to be showing. That last part is the fix for a real bug: the transcript panel (way 2 below) has no language field on it, it simply answers in whichever language the panel opens on, which is the caption track you last switched on by hand. Asked blind, it handed back the English transcript of an Italian video, and Torval, having asked for a transcript and been given one, used it. It is now asked for a named language, through the panel's own language menu, and a transcript that cannot be shown to be in the right language is refused rather than used, which costs one request and never costs correctness. The address caught in way 1 is checked the same way, against the lang in it, so an English track the player fetched for its own reasons no longer overwrites a correct transcript.

Torval leaves the CC button alone. It is YouTube's and means what it says, Torval's line comes from the track it fetched itself, and either can be on without the other. Both at once is then a choice rather than an accident. The one exception is the last of the four ways of getting the timing below, which works by reading YouTube's captions off the screen and so cannot work with nothing on the screen to read: there, once everything else has failed, Torval turns the player's own captions on, in the language being read, using the same call the player's own settings menu makes. Only on, never off, and only to that one language. Before, this case left nothing on screen and said so only in the console.

Netflix

There are two ways in, and the one that reads better on paper is not the one that works.

What works: taking a copy of the file as it goes past. Netflix's subtitles are not part of what the DRM protects. The player downloads them from oca.nflxvideo.net as an opaque ?o= address with no file extension and TTML inside it, in the clear, like any other file on any other site. So Torval watches what the page fetches, and anything whose body turns out to be WebVTT or TTML is the subtitle file, caught whole, with every timing in it. No manifest, no injected format, nothing to guess.

The file is only downloaded when the player is going to show something, which would mean going into Netflix's own menu, turning on the language you are reading, and then watching two sets of subtitles at once. The player will do it when asked, though: it keeps a list of its tracks and a method to choose one, the same pair its own menu is built on. So Torval turns the track on itself, waits for the file that fetches, and puts your own choice straight back. What is left behind is the player exactly as it was and the whole subtitle file in hand. A moment of Netflix's own subtitles may flash up in between, once per episode. If you had already turned that language on yourself, nothing is touched at all.

This is what asbplayer does now as well. It used to carry Netflix-specific code and today has not one file with Netflix in its name: it watches replies for anything shaped like subtitles instead. A site that rearranges its internals every few months cannot be followed by knowing its internals.

The other way, kept because when it works it is better: the whole track list before a second has played, which sounds unlikely and turns out to be a matter of asking.

When the player starts a title it posts a request listing the formats it is prepared to accept, and the answer only ever offers what was asked for. The player never asks for WebVTT, so the answer never offers it, and there is nothing to find afterwards however hard you look. Add that one format to the request on its way out and the very same answer comes back with a plain, unencrypted address for every subtitle track in it, Japanese included.

So Torval adds it. JSON.stringify is where the request becomes text, which is the last moment before it is sent, so that is where the format goes in. The track list that comes back is offered to the extension rather than filtered on the page: page code knows nothing about which language Torval is set to read, so it hands over every track the title has and is told which one to fetch back.

On the account this was written against, the request goes out and its answer is read somewhere none of those hooks can see. That is why the file is caught on its way to the player instead, and why both are kept: this one costs nothing when it fails, and hands over the whole track list when it does not.

Reading the answer took longer to place. JSON.parse is the obvious spot and it is the wrong one: nothing carrying a track list ever went through it, because response.json() does not call JSON.parse at all, the browser parses the body itself. So the answer is watched at all three places a reply can become an object, JSON.parse, Response.prototype.json and an XMLHttpRequest’s responseText, and at the last two only for a reply to the manifest request, so nothing else on the site is touched.

Netflix’s player is not affected by any of it: the extra format is one more line in a list it ignores, and every answer is handed straight back, the same object and the same promise the real one produced.

Getting the hooks onto the page took four goes

They have to run as page code. Netflix’s JSON is not an extension’s JSON, and every part of that site parses JSON for everything it does, so anything less than being page code shows up as the site not working. Three ways failed first, and they are worth writing down:

  1. A "world": "MAIN" entry in the manifest’s content_scripts never ran at all, on a Firefox new enough to support the property. Most likely Firefox rejects the whole entry over it rather than ignoring it.
  2. A <script> tag added by a content script never ran either, near certainly stopped by Netflix’s content security policy.
  3. Reaching the page’s own JSON from the content script, through wrappedJSObject and exportFunction, did run, and broke Netflix: the home page came up with its header and nothing else. Passing every parsed answer on the site across the wall between two worlds is not free.

What works is registering the file from the background script through the scripting API, with the same world: "MAIN" the manifest would not take. It follows the switch on the toolbar button, so turning Torval off takes it off Netflix altogether: after the third attempt above, a way out that is not "uninstall the extension" seemed worth having.

When there is no file

For a title whose file never turns up, or an extension loaded halfway through an episode, Torval reads the lines off the screen as it does on YouTube, and the language being read has to be the subtitle language turned on in the player. It is a poor second: the percentage can only describe the lines already watched, and D has nowhere to go, because the next line has not been said yet. The console says which of the two is in use.

A and D ask Netflix’s own player to move rather than setting currentTime on the video element. Netflix streams in pieces chosen in advance, and moving the element under it ends the session with error F7375 and an error page.

One thing is different from YouTube either way: on Firefox the audio on a mined card does not record, because the browser has no way to hand an extension a tab's sound, and a decrypting video element hands over none of its own. On Chrome it does, once the box under Settings → Anki is ticked. Either way the card is made, with the sentence and the word on it, and it says which half is missing. See the capture section further down.

There are four ways it gets the timing, tried in the order below, each a fallback for the one before it, not a choice between them. The first three all build a caption web address themselves, out of data YouTube's own page publishes; none of them ever produced anything, on any video tried, no matter how the request was made, a background-script fetch, a content-script fetch, a rewritten request header, an injected response header. What finally worked was not sending a better-formed request at all: it was not building the address in the first place.

  1. Catch the address YouTube's own player already uses. The address published in the page's own data is not, evidently, the one YouTube's own player actually requests when it genuinely fetches a caption track, so no amount of asking more carefully for that address was ever going to work. Torval watches the page's own network traffic for the real request instead, which happens the moment a caption track is genuinely active in the native player, and reuses that exact address. This needs a real caption track active at least once. The address says which language it is for, and one in any other language is ignored.

    Catching the right address turned out not to be enough on its own. Even that address, provably the one YouTube's own player had just used successfully, still came back with a 200 and nothing in it when refetched from the content script, the same "blocked by OpaqueResponseBlocking" symptom from the very first attempt, which meant it was never really about which address was being asked for. Both fixes are needed together: Torval also adds the CORS permission the response never carries, using the same webRequest technique CORS-unblocking extensions use generally, scoped only to this one address.

  2. Ask for the transcript the way "Show transcript" does. Not the closed-caption file, the separate panel YouTube's own player offers, reached through a one-time token buried in the page's own data.

  3. Ask YouTube for the closed-caption file directly. The address Torval builds itself from the page's published data, tried as a fallback in case a future change makes it work again.

  4. Read the captions off the screen as they play, timing each line by watching it appear and disappear. This is what the simplest subtitle tools do, and it always works, because it is only reading what is already there. The real cost: a line is known only once it has actually been shown, so nothing about the video is known ahead of watching it, D cannot jump ahead into an unseen line, and nothing here can answer "how much of this video will I understand" before you have already watched it. Rewatching a line does not duplicate it; seeing the same text again near where it was last seen just refreshes its timing.

Method 1 can arrive at any moment and supersede whichever of the others is currently in charge, including the on-screen fallback, since a genuine transcript beats one assembled a line at a time regardless of when it turns up.

The line can be moved up or down the picture, when it sits over something worth seeing: press on it and pull. Where you leave it is remembered, so a video watched tomorrow puts it back there.

Moving it and selecting the words in it are one gesture and had to be told apart, and the direction does it, because the two never really point the same way. A subtitle is wide and one line tall: selecting it means going along it, moving it out of the way means going up or down. So the first few pixels of a press decide which one it is, and it stays that until you let go, a clear pull up or down moves the line, anything else selects as it would on any other text and the line does not budge. Before, any press that moved at all moved the line, so copying a word out of a subtitle shoved it up the screen.

Whichever way found the timing, A steps back a line and D forward. Part-way through a line, A restarts it; pressing it again goes to the line before, which is how you rewatch something you did not catch.

A manually authored track is always preferred when one exists, since auto-generated captions are also where the word-by-word reveal mentioned below comes from.

In the on-screen fallback specifically, auto-generated captions are often revealed a few words at a time as recognition catches up rather than appearing whole. Torval treats a growing or slightly revised line as the same line still being written, not a new one each time, otherwise both the audio and the timing would start wherever the last fragment happened to begin, not at the sentence's true start.

Pressing + on a word in a subtitle also puts on the card the frame you were looking at and the audio of that line. The frame is taken the instant you press it, before anything moves. The audio is taken by replaying the line: the video is sent back to the start of it, recorded to the end of it, and put back exactly as it was, same moment, same speed, same paused or playing.

The line plays out loud while this happens, so you hear what is being put on the card rather than mining blind. It takes as long as the line does, a couple of seconds, and what comes out is exactly the line.

When it finishes, the video is left where the recording ended rather than dragged back to where + was pressed: you have just heard the line played out, and the place to carry on from is the end of it. Never earlier than where you were, though, so pressing + on a line that has already finished does not rewind you into it.

A second + pressed while one recording is running waits for it rather than being turned away. Both cards get their sound; the second takes a few seconds longer.

Only a subtitle brings the video with it

A frame and a line of audio belong to a word read off a subtitle. Mining one out of a comment under the video used to put whatever happened to be playing on the card, for two reasons at once: the frame was grabbed whenever there was a video anywhere on the page, and the line to record is found by matching the sentence against the subtitles, which falls back to the line playing now when nothing matches. That is a fair guess about a subtitle and nonsense about a comment.

The sentence now remembers whether it came off a subtitle, and nothing reaches for the video unless it did.

A line is not always one cue

An automatic caption revises itself as the recogniser hears more, and every revision is filed as a cue of its own. What you read as one line is several cues in a row, each a rewrite of the one before, and any single one of them can be well under a second long.

Mining used to pick one of those and record exactly it, which produced a clip of the lead-in and nothing else: half a second of the previous line, stopping at the moment the line you wanted began. A line now runs from the first of those cues to the last, and a clip is never shorter than 1.2 seconds whatever the timings say. On a clean subtitle track, where each line is filed once, neither rule changes anything.

Playback speed is forced to normal while it records, since a line captured at 1.5× is a line spoken at 1.5×.

The clip goes on the card as a plain WAV, mono and 24 kHz, not as what the browser recorded. A browser records Opus in a WebM container, and Anki on a computer plays that happily because it hands the file to mpv, which plays anything. Anki on a phone does not: AnkiMobile cannot read WebM at all, and AnkiDroid depends on what the phone underneath it supports. A card that plays at the desk and is silent on the train is worse than useless, because the train is where you find out.

The cost is the file: a couple of hundred kilobytes for a line where the Opus was twenty or thirty. There is no honest way around that without shipping an MP3 encoder, which is a large piece of somebody else’s code for a problem that only exists on a phone. 24 kHz carries everything a voice does and keeps it in proportion.

Content-protected video, and what that actually stops. Netflix, Prime Video and Disney+ hand their video to the browser's DRM layer. The two halves of mining fail there for different reasons and only one of them is fixed, which is worth separating, because "DRM, nothing to be done" is both the easy answer and the wrong one. Other mining tools do get pictures off Netflix.

The picture can be had, the long way round. A protected frame does not refuse to be drawn. It draws as one flat black rectangle: nothing throws, the canvas is not tainted, and what would go on the card is a black JPEG, which is worse than no picture because it looks like a card that worked. So Torval takes the frame, then looks at it: a frame that is one single value across its whole area is not a frame, and is thrown away rather than put on a card. A real picture is never that flat, however dark, and the one thing that is, a deliberate fade to black, is a cheap thing to be wrong about.

Turning the browser's hardware acceleration off used to be the answer, and on Firefox it still is: the frame stops living in a GPU texture on the protected path and draws as itself. On Chrome the setting is no answer at all on plenty of machines, because the refusal is not really about where the frame was decoded. The element will not hand its pixels to page script, and page script is what a canvas is.

So the frame is taken from outside the page instead. tabs.captureVisibleTab is the browser photographing what is on screen, the picture a person is already looking at, and by the time it exists the protection has had its say. Torval crops the video out of that photograph using the element's own rectangle, scaling by the ratio between the image and innerWidth rather than by devicePixelRatio, which is wrong the moment a page is zoomed. Two things follow. The site's own subtitles are drawn over the video and would land on the front of the card, so they are hidden for the one frame and put straight back. And this is the top window only: the coordinates belong to it, and a video in a frame is a YouTube embed, which is not protected and never takes this path.

Both the photograph and the tab recording need the extension to have been invoked on the tab, which is Chrome's activeTab rule and has no way round it. One click on the toolbar button covers the tab until it navigates. When the browser refuses for that reason the card says so in those words, because it is the one failure here with a cure the reader can apply.

The sound needs a different mechanism, and one browser has it. captureStream() on a decrypting element hands over no audio track at all, so the ordinary path has nothing to record. The tab's own output is a different thing entirely, just sound coming out of a tab, and Chrome hands that over through tabCapture. That is what Torval does there, and what asbplayer and Migaku do.

Firefox cannot, and this is the one thing in Torval that works in one browser and not the other. It has no tabCapture API at all, and its getDisplayMedia ignores the audio option without an error or a warning, which is Mozilla bug 1541425, filed in 2019 and still open. There is no third mechanism and nothing to fall back on. The Firefox build ships without any of it: package.mjs takes both manifest keys and both offscreen files out rather than leave a reviewer to work out why Chrome-only code is in the package, and the settings panel says plainly that this one needs Chrome.

On Chrome it is off until asked for. tabCapture is an optional permission, requested at the moment somebody ticks the box under Settings → Anki → Recording from video and at no other time: not on install, and not by watching a video. It is used on a copy-protected video and nowhere else, because YouTube's sound comes off the element and always has.

Two details that are not optional. The recorder lives in an offscreen document, because a Manifest V3 service worker has neither getUserMedia nor MediaRecorder. And capturing a tab takes its sound away from the speakers, so the captured stream is played straight back out through an AudioContext; without that the video goes mute mid-sentence, which looks exactly like a crash, and mining blind is the thing this design exists to avoid. The seek to the start of the line goes through the site's own player, the same way A and D do, for the same F7375 reason.

Whichever half is missing, the card is still made, and the popup says which and why, and sends anybody who wants the longer answer to Settings → Anki. The one exception is the uninvoked tab, which gets its own sentence naming the toolbar button, since that is a cure rather than an explanation. Torval asks the video element whether it is decrypting rather than keeping a list of sites: the page sets mediaKeys on it when it starts, which stays right on the day a fourth service launches and on the day one of these three plays an unprotected trailer.


Known words

It is off until you ask for it

Torval does two things and only one of them is a dictionary. The other is the counting: marking a word known or ignored, colouring what is left, and the percentage across the top saying how much of a page is built from words you have. That half is what makes Torval more than a dictionary, and it is also the half that asks something of you first. A word list starts empty, so the number is wrong until a few hundred words are in it, and somebody who only wanted to know what a word means has been handed a chore they did not ask for and a percentage that lies to them for a fortnight.

So it is off on a fresh install, and the switch is the first thing on the Words panel. With it off Torval is a pop-up dictionary that makes Anki cards: hover, read, press +. No bar across the page, no colours, no ✓ and ⊘ in the popup, those keys do nothing, no page is read at all, and no daily copy of a list that does not exist. Nothing to set up and nothing to explain.

It is never turned off under anybody. An install that already has words in its list has answered this question already, so wakeTracking in background.js looks at the counts once on the way past and turns it on; after that the switch is the only thing that moves it, including back off again. The answer lives in storage rather than being passed around, because a content script, the background and the settings page all need it at once and storage is the only thing all three share. See track.js.

The lists themselves

The list of words you already know, kept under Words in Torval's settings, alongside the ignored list: they are the same decision with different answers, and a word moves between them, so they sit on one page. Words get on the known list two ways.

One at a time. Every word in the popup has a ✓ beside its +. They are different questions: + means teach me this, ✓ means I already have this. The tick toggles, because the commonest mistake to make with it is pressing it on the wrong word.

Or press 2 while pointing at the word, without reaching for the tick. Most of what you meet while reading is something you already know, and saying so is the one thing done often enough that it should not cost a mouse movement.

The three keys run in the order the answers themselves run: 1 for a word you do not know, 2 for one you do, 3 for one to stop mentioning. They set rather than toggle, so leaning on one is harmless, and every state can be reached from every other: 1 on a word already marked known puts it back to unknown, which used to mean a trip to the popup.

In bulk. Paste in something you have already read, or load a plain text file, and press Add words from this text.

Nothing new is built to read that text: it runs through extractWords, the exact longest-match search a Shift-hover already uses, moved forward across a whole passage instead of stopping at the first word. 走っていました is recorded as 走る, the same dictionary form a hover on it would show, so reading a passage once teaches the word regardless of which sentence it turned up conjugated in, and running the same text through a second time adds nothing, since it is already known.

The same page browses the list, newest first, with a search box and an × per word for the ones added by mistake.


Reading a book

Open your books, at the foot of Torval's settings, opens a page that takes an epub or a txt file and shows it as an ordinary web page. That is the whole trick: the reader loads the same scripts Torval puts on any website, so hovering, the popup, the marking and the comprehension bar work on a book exactly as they do on a page, with no second copy of anything.

An epub is a zip of XHTML files plus a list saying what order to read them in. Torval reads that list, takes the text out of each file, and shows one chapter at a time, named by the book's own table of contents. Most books do not repeat the chapter name inside the chapter, so without reading the contents a book opens as Section 1, Section 2, Section 3, which is no way to find your place. EPUB 3 keeps that list in a nav document and EPUB 2 in a toc.ncx, and books in the wild are still mostly the second kind.

Furigana is thrown away rather than kept: left in, every word would arrive with its reading glued to it and 食べる would come out as 食た べる, matching nothing. Text is taken from the innermost blocks only, since a paragraph inside a blockquote sits in two of them and would otherwise be read, shown and counted twice. Books built out of bare <div>s or nothing but line breaks are common enough to be handled the same way as ordinary paragraphs. A chapter longer than about twelve thousand characters is cut into parts, both because that is a long way to scroll and because reading a page end to end for a score should take a moment rather than a minute.

The reader opens on a shelf: every book you have added, with how much of each one you would understand beside it, coloured the same way the bar is. That estimate comes from a sample taken evenly through the book rather than from the whole of it, because the whole of a novel is a minute of reading and the answer would not move. Books fill in one at a time as they are worked out.

The bar stays down inside the reader instead of tucking itself away, since a page that exists only to be read in has room for it, and hiding the one number you are there for would be strange.

Books and your place in each are kept in the browser's storage, so closing the tab loses neither. Your place is the chapter and how far down it you had read, since a chapter is several screens.


Comprehension

A slim bar across the top of any page says how much of what is in front of you is made of words you already know. By default it stays out of the way: a small "Torval" handle sits in the top right corner, and clicking it slides the bar down.

Torval   87%   1,204 of 1,383 words known                    ⟳  ⚙  📌  ×

Left unpinned, it is a hover panel: moving the mouse away tucks it back up after a moment, so it never sits permanently across the top of a page you did not ask it to. Press 📌 to keep it down instead, the way it used to work; that choice is remembered.

On a video that is measured against the whole transcript, not the part already watched, which is the entire point of the fight to get the transcript up front. Knowing a video is 87% words you know before starting it is what decides whether it is worth watching; reading it afterwards answers nothing. Everywhere else it is the page's own text.

It counts every word said, not every distinct word. A page that says 私 forty times and one word you have never met is not as hard as one with forty different unknown words in it, and a score that could not tell those apart would not be worth reading.

Pressing ✓ in the popup moves the number immediately. A word's count is exactly how far the bar shifts, so nothing has to be read a second time. ⟳ reads the page again, ⚙ opens the settings, 📌 pins it open. It hides itself while a video is full screen, and never appears at all on a page with no Japanese on it.

While the dictionary is still being built, or a page is still being read, the handle says so rather than sitting there silently: "Torval 42%" the first time, when 218,000 entries are being copied into the browser's own database, and "Torval ·" for the moment a page takes to read.

A grammatical pattern JMdict happens to file as one entry, お元気ですか ("how are you") is nothing more than the honorific お, 元気, the copula です and the particle か, counts as known once every piece of it is, even though that exact four-word entry was never separately marked known itself. The same goes for ことがある. A true idiom, where the meaning genuinely is not the sum of its words, gets none of this: JMdict's own "id" tag is what tells the two apart, so 猫の手も借りたい stays unknown no matter how well you know 猫 and 手.

Any reading of what is actually written counts, not only the best one. 来た is the past tense of 来る and also, on paper, a rare interjection; 読み is the stem of 読む and also a noun in its own right. Knowing either reading of what is on the page means nothing is missing, so an i-stem never counts as a new word just because the dictionary also lists it separately.

Ignored words

Some words are never going to be learned and should not be counted either way: names, pieces of English, things the dictionary read wrongly. Press ⊘ in the popup, or the 3 key, and the word leaves the question entirely, not marked on the page, and out of the comprehension total rather than counting for or against it. Counting them unknown would say a page is harder than it is; counting them known would say the opposite; neither is true.

Known and ignored are the same kind of decision with different answers, so a word is one or the other or neither, never both: putting it on one list takes it off the other. Both are browsable under Torval's settings, and either can be taken back.

Does counting words this way actually make sense?

It is the standard approach every tool like this uses: percentage of running words already known, counted once per time a word is actually said rather than once per distinct word, so a page that says 私 forty times reads differently from one with forty different unknown words in it. That much is sound.

Two honest limits are worth knowing about. First, this is a vocabulary score, not a comprehension score in the full sense: understanding a sentence also takes grammar, and two sentences with the same known-word percentage are not always equally easy to follow. Second, a word the dictionary does not recognize at all, a name, a coined word, a typo, is left out of the count entirely rather than counted as unknown, since there is no dictionary entry it could be. On a video full of character names this reads a little higher than it should. Ignoring a word does the same thing deliberately, and for the same reason; the difference is that this happens without being asked.


Keeping a copy of your words

At the end of the Words page in Torval's settings, Save to a file writes both lists, known and ignored, into a single JSON file with the date in its name.

This is the one part of Torval that cannot be rebuilt. The dictionary downloads again in a minute and the Anki settings are a minute of typing, but a known list is however many months of reading. Torval keeps a copy for you as well, once a day, described below; this is the button for when you want one now, or want it somewhere of your own choosing.

Loading a file back adds to what is already there. Nothing is removed and nothing is overwritten, and where the same word is in both, the earlier of the two dates is the one kept. That makes an old backup safe to restore: it can only ever give words back, never take away ones learned since, so carrying one file between two machines works in either direction. If a word is known on one side and ignored on the other, known wins, since the two lists still cannot both hold it.

A file is read, not copied in. Torval's own files hold dictionary forms already and come back as themselves, but a file written by something else holds whatever was in the field it was told to read: a whole sentence on a sentence deck, a speaker's name, a line of English on the back of a note. So every key goes through exactly the reading a page, a subtitle line or the Add from a text box gets, segmented and deinflected and looked up, and what comes out is the dictionary forms found inside it. A sentence becomes its words and an inflected hablaba becomes hablar.

What the dictionary does not recognise does not go on the list. The known list is what a page is measured against, and a word that can never be met again while reading cannot do anything there except inflate the number; a whole deck in the wrong language would inflate it enormously. The count of what was left out is shown beside the count of what went in, so a file that added nothing says why.

The ignored list is the exception, and is taken exactly as it comes. Those are the words the dictionary has nothing for, which is the reason they are on that list. Asking it to confirm them would throw away precisely the list.

A file that is not one Torval wrote is refused rather than half-read. Files written before this was called Torval say lll-words inside and are read too: a word list is the one thing here that cannot be rebuilt, and refusing last month's copy of it because the program has since been given a different name would be the worst possible reason to lose one. They were saved to Downloads/LLL; new ones go to Downloads/Torval, and both load back the same way.

Starting over

Under Starting over, at the very bottom of the same page, is a button per list that empties it. It is there for the one thing nothing else undoes: words in the wrong language, or somebody else's list loaded by mistake, or an Anki deck that turned out to be the wrong deck.

Each button has to be pressed twice. The first press only changes it to Press again to forget them and starts an eight-sec