Lexical Coverage (Spanish only)

AnkiWeb addon 861555972

Analyzes tagged Spanish cards to estimate CEFR vocabulary level, corpus coverage, and band breakdown using frequency lists.
AI-generated summary; may contain mistakes.

statisticsprogress-tracking

Open on AnkiWeb GitHub Ask about alternatives

AnkiWeb

Rating
0 (πŸ‘ 0 Β· πŸ‘Ž 0)
Updated
2026-06-18
Anki versions
2.1.1~
Description language
en

Maintenance

active

  • Last update or commit was 105 days before the snapshot (2026-06-18).

Will it work on my Anki?

Version branches
Min AnkiMax AnkiUpdated
2.1.1+2026-06-18
History across monthly snapshots

Loading…

Similar addons

Tag Ratio Panel

Calculates tag-based coverage and completion ratios across selected Anki decks, with customizable tag logic, deck scopes, color bands, and optional auto-update.

statisticsprogress-tracking

Rating 0 πŸ‘ 0 Β· πŸ‘Ž 0 ⭐ 0 Anki 23.10~ Updated 2025-12-22
JPDB Mining Progress

Scans configured Anki note fields, matches mined Japanese words to the JPDB top 10k by word and reading, and displays coverage, ranks, searchable lists, unmined entries, and manual marks.

statisticsprogress-trackingdictionary

Rating 1 πŸ‘ 1 Β· πŸ‘Ž 0 Anki 25.06~ Updated 2026-06-20

README

Lexical Coverage Analyzer

An Anki add-on that estimates your Spanish vocabulary level by analyzing the words on your cards and comparing them against a frequency-ranked word list. It reports your estimated CEFR level (A1–C2), corpus coverage percentage, and a breakdown of your known vocabulary by band.


Requirements

  • Anki 2.1.50 or later
  • Internet connection on first run (to install dependencies)

The add-on will automatically install the following Python packages on first run:

  • spaCy β€” for lemmatization of card text
  • matplotlib β€” for the coverage chart
  • es_core_news_sm β€” spaCy's small Spanish language model (~12MB download)

Installation

  1. Download the lexical_coverage folder
  2. Place it in your Anki add-ons folder:
    • Mac: ~/Library/Application Support/Anki2/addons21/
    • Windows: %APPDATA%\Anki2\addons21\
    • Linux: ~/.local/share/Anki2/addons21/
  3. Restart Anki
  4. On first launch, Anki will install the required dependencies β€” this may take 30–60 seconds depending on your internet connection

Setup β€” Tagging Your Cards

The analyzer works by analyzing all cards with a specific tag. Before using it, make sure all your Spanish cards are tagged consistently. We recommend using a single tag like Spanish for all your Spanish vocabulary cards.

To tag cards in bulk:

  1. Open the Anki Card Browser (Browse in the main window)
  2. Select the cards you want to tag
  3. Right-click β†’ Add Tags and enter your tag (e.g. Spanish)

Usage

  1. Open Anki and go to Tools β†’ Lexical Coverage Analyzer
  2. Enter the tag you used for your Spanish cards (e.g. Spanish)
  3. Click Analyze
  4. Results will appear after a few seconds

Understanding Your Results

Estimated CEFR Level

Your overall vocabulary level based on how much of the most frequent Spanish words you know. Levels range from A1 (beginner) to C2 (near-native).

Level Words covered Coverage threshold
A1 Top 1,000 90%
A2 Top 2,000 92%
B1 Top 3,000 94%
B2 Top 5,000 96%
C1 Top 8,000 98%
C2 Top 25,000 99%

Corpus Coverage %

The percentage of words in real Spanish text that you would recognize. Based on subtitle corpus frequency data:

Coverage What it means
~90% You understand most casual text but frequently hit unknown words
~95% Comfortable reading with occasional lookups (B2 territory)
~98% Near-fluent reading comprehension (C1+)
~99% You almost never encounter unknown words

CEFR Band Chart

A bar chart showing your coverage percentage for each CEFR band. Green bars indicate bands you have reached; red bars indicate bands you haven't reached yet. The dashed line on each bar shows the target threshold for that level.

Word Lists

Expandable lists showing your known words sorted by CEFR band, plus words beyond C2 and words not found in the frequency list. Click any section header to expand it.


Output Files

The add-on writes two files to the lexical_coverage/ folder after each analysis:

  • unknown_words.txt β€” a full list of words from your cards that weren't found in the frequency list. Use this to spot misspellings in your cards or identify vocabulary gaps.

Limitations

  • Card quality matters β€” misspellings and typos in your cards will appear as unknown words. Use unknown_words.txt to identify and fix these.
  • Proper nouns β€” names of people, places, and fictional characters will always appear as unknown words since they're not in the frequency list. This is expected.
  • Non-Spanish words β€” if your cards contain English translations or other languages, those words will appear as unknown.
  • First field only β€” the analyzer reads only the first field of each card (conventionally the front). Cards with non-standard note types may not be analyzed correctly.
  • Analysis time β€” analyzing a large deck (2,000+ cards) may take several seconds. Anki's UI may be unresponsive during this time.

Frequency List

The add-on uses a lemmatized Spanish frequency list derived from the Hermit Dave FrequencyWords project (Creative Commons licensed), covering the 50,000 most common Spanish words from subtitle corpora. The list has been lemmatized using spaCy's large Spanish model (es_core_news_lg) to ensure consistent matching against your card text.


Reporting Issues

If you notice words being incorrectly lemmatized or other unexpected behavior, please open an issue on the project repository. Including a sample from your unknown_words.txt is helpful for diagnosing lemmatization problems.