Log Your Mistakes, Let AI Sort Them, Push Them to Anki

You found the holes with mutual explanation, checked them with generated questions, and worked through a problem set. Then, two weeks out, you redo the official samples and the error you already fixed, that pseudonymised information may be provided to third parties, is back. Knowledge you learned disappears unless it is recalled at intervals. And there is no time to run through everything again.
This article contains affiliate links. As an Amazon Associate, we earn from qualifying purchases.
This article is about switching G-test memorisation from run through everything to run through only your mistakes. Plenty of people use Anki for the G-test; getting an AI to help decide what becomes a card cuts the time spent making cards and skews the cards you do make towards your weak points. Collect the mistakes from mutual explanation, problem sets, generated questions and mock exams into one table, have the AI aggregate it by major item and error type, and turn only the weak terms into cards for spaced repetition. This site built ten cards from a 21-row mistake log and imported them into Anki; the results are below. Figures and specifications are from official material as of 21 September 2026.
Why run only the mistakes
The final fortnight of a 40-day run gives you about ten hours at 45 minutes a day. Reviewing all 495 syllabus keywords, counted per mid-level item, leaves a little over a minute each. And most of them are words you can already state. Most of the time goes to reconfirming what you know.
Running only the mistakes inverts that. The words you could not say and the pairs you got wrong are already in the record. In our measurement there were about 20 rows by the middle of the 40 days. Twenty rows fit inside ten hours even at 30 minutes each.
The other reason is the shape of forgetting. An error you have just corrected is available immediately afterwards and tends to revert within days. Recalling it tomorrow, then in three days, then in a week, fixes it in fewer repetitions than seeing it at the same interval every day. Managing when to surface each item by hand is tedious, and software such as Anki does it. The general literature on learning is out of scope here; this is procedure and measurement.
The mistake log: a four-column table
Every mistake, whatever stage produced it, goes into the same table. Four columns are enough.
| Column | Content | Example |
|---|---|---|
| Origin | Which stage produced it | Mutual explanation, problem set, generated question, mock exam |
| Major item | The syllabus major item name | Law and contracts around AI |
| Term | The term or the pair you got wrong | Anonymised versus pseudonymised information |
| Error type | Definition swap, wrong attribution, missing condition | Missing condition (the rule and its exception) |
Decide per origin what goes in the note field. For an error flagged in mutual explanation, copy the substance of what the AI said (zero covariance means uncorrelated, not independent). For a wrong answer in a problem set, write the question number and the substance of the option you picked. For something found while verifying a generated question, write the name of the correct primary source. For a mock exam the question number is enough, but adding a separate row with the accuracy rate per major item after the mock lets the aggregation show you both which major item is weak and which term is weak.
Use the syllabus spelling for major item names. If they drift, the aggregation will count machine learning and overview of machine learning as separate rows. One reason the glossary from the earlier step carried mid-level item numbers was to align the axis of this table with the structure of the exam.
In our measurement, seven planted errors from the mutual-explanation article plus fourteen defects found while verifying generated questions (across thirteen questions; one question had two kinds of defect) made a 21-row table. The second group are the AI mistakes, but they are terms the learner had to go and look up, so they count as the learner weak points too.
Let the AI classify: aggregation and priority
Twenty-one rows can be aggregated by hand, but handing them to an AI produces the aggregation, the priority order and draft cards in one pass. The instruction we used:
Below is a learner mistake log in CSV.
(1) Count the rows by major item crossed with error type.
(2) List the top 10 terms to review first.
(3) For each, write an Anki front (the term) and back
(a one-sentence definition plus how it differs from the term it is confused with).
----- CSV -----
(paste the mistake log)The output was a 16-cell aggregation, ten priority terms and ten draft cards. Checking all 16 cells against the original table, every count matched, totalling 21. That is one run over 21 rows, so accuracy at larger row counts is untested. Keep the habit of reconciling the aggregate total against the number of rows.
You can leave the prioritising to the AI, but it is better to state the criteria. We left them unstated, and the seven pairs planted in mutual explanation came out on top, with the rest being proper nouns from law and guidelines found during verification. Next time we will add that a term appearing from more than one origin ranks higher, and that definition swaps rank above missing conditions. Once mock-exam mistakes join the log, put those at the very top: what you got wrong in the format closest to the exam, closest to the exam date, is what matters most.
Do not use the draft cards as written. Check the definition on the back against a primary source first. Of our ten, two needed changes. One was the anonymised versus pseudonymised information card, whose back described the condition for third-party provision of anonymised information as publishing the processing method, when under article 43 of the Act on the Protection of Personal Information the processing method is information to be kept secure; what is published is the categories of information contained and the method of provision. The same card described the exception for third-party provision of pseudonymised information as entrusted processing and the like. The exception the statute sets out is where required by law (articles 41(6) and 42(1)); entrusted processing, business succession and joint use are handled instead by the application of article 27(5), under which the recipient is not a third party at all. Whether something is an exception or falls outside the prohibition is a difference in where it sits in the statute. The other card was on the contract guideline on the use of AI and data, tagged under social implementation, when the syllabus places that term in the legal domain at mid-level item 6, AI development outsourcing contracts. Legal cards go wrong on a single word between the rule and the exception, so run them past the statute or an official explanation once.
Having the AI classify a mistake log is a structured-output prompt. For techniques that make classification prompts consistent and easy to check, this book is a thorough guide.
What Anki is
Anki is flashcard software for spaced repetition. The desktop version (Windows, macOS, Linux) and syncing through AnkiWeb are free. AnkiDroid on Android is also free and is a separate project maintained by contributors. AnkiMobile on iOS is paid; check the price on the App Store. The current version as of 21 September 2026 was 26.09.2.
The loop is: look at the front, recall, turn it over, grade yourself, and the software decides when to show that card again. The division of labour is the AI making cards, Anki scheduling them, you answering them.
A spreadsheet or a vocabulary app can do something similar. Three things make Anki the choice here. Importing from a text file is stable when you use header directives, so an AI-produced CSV flows straight in. Tags are hierarchical, so you can review one major item at a time. And the desktop version and syncing are free, so 40 days of study incurs no new cost. Spaced repetition itself exists elsewhere; those three points are what mesh with pouring in an AI-built table every week.
Building the CSV: header directives make imports stable
Anki imports cards from a text (CSV) file. Put directive lines beginning with a hash at the top and you skip choosing the separator, note type, deck and tags in the import dialog every time (Anki 2.1.54 and later). The head of the CSV we generated:
#separator:Comma
#html:false
#notetype:Basic
#deck:G-test
#tags:G-test weakness
#columns:Front,Back,Tags
Supervised vs unsupervised learning,Supervised learning learns an input-output relation from labelled data; unsupervised learning finds structure in unlabelled data. Note that clustering belongs to unsupervised learning.,G-test::Overview of machine learningEach row is one card: first column the front, second the back, third the tags (the back in the example is abbreviated). When the back contains the separator, a line break or a quotation mark, wrap that whole field in quotation marks; to include a quotation mark itself, double it. For line breaks, set html to true and use a br tag.
Putting the major item into the third column as a hierarchical tag lets you filter review by major item inside Anki. Use the major item column of your mistake log as the tag.
Measured: importing ten cards, and re-importing without duplicates
To check that the CSV actually imports, this site used the official Anki Python library (version 26.09.2, the same number as the current desktop release). Creating a fresh collection and importing the CSV above gave:
| Item | Result |
|---|---|
| Header interpretation | Separator comma, HTML disabled, tags G-test and weakness, note type Basic, deck G-test, all recognised |
| First import | 10 notes found, 10 new, 0 updated, 0 duplicates. 10 notes and 10 cards. The G-test deck was created and the major item tags applied |
| Re-importing the same CSV | 0 new, 0 updated. The note count stayed at 10 |
The ten imported cards broke down as two on the overview of deep learning, two on mathematics and statistics, two on law, and one each on the overview of machine learning, what artificial intelligence is, social implementation and ethics. Seven were pairs planted in mutual explanation; three were proper nouns from guidelines found while verifying generated questions. One of the two legal cards had carried the major item straight over from the mistake log and arrived tagged social implementation. Copying the major item column into the tag is fast, but if the log has the wrong major item so does the tag. Check the tag against the syllabus table once before importing. One card, as an example: front, variance versus covariance; back, variance is the spread of one variable and covariance the tendency of two variables to move together, and note that zero covariance, meaning uncorrelated, does not imply statistical independence; tag, G-test::Mathematics and statistics for AI. The error planted in mutual explanation became the card easily-confused-with line.
The third row is the important one. By default Anki treats a note whose first field matches an existing note as a duplicate and updates it if the content differs, doing nothing if it is identical; the setting is called Existing notes and its default value is Update. That is, you can refresh the mistake log weekly, rebuild the CSV and import into the same deck without the card count growing, and the scheduling on the existing cards is preserved. That is what makes a weekly cycle viable. Two caveats: matching is scoped by note type through a separate Match scope setting, and because matching keys on the first field, editing the front in your CSV produces a new note rather than an update.
Note that this check was done through the library, not through the desktop import dialog. The dialog should read the header directives the same way, but confirm on the first run that the deck and the tags came out as intended.
The weekly cycle for the final fortnight
| When | What | Time |
|---|---|---|
| Daily | Anki review (answer the cards it surfaces) | 10 minutes |
| Daily | Solve a few generated or printed questions and append mistakes to the log | 20 minutes |
| Weekend | Hand the log to the AI for aggregation and new card drafts. Verify against primary sources, add to the CSV, import | 30 minutes |
| Final week | Run one mock exam and append the mistakes. Card them in the same weekend routine | 100 minutes plus 30 |
Keep new cards to ten or twenty a week. Anki has a per-deck limit on how many new cards are introduced each day, called New Cards/Day under Daily Limits; set it to about five. Beyond that, the review count balloons within days and no longer fits in ten minutes. Rather than carding every row of the log, take the top of the priority order the AI produced.
Do not add new cards the day before the exam. Run only what comes up for review and look again the next morning at whatever you could not answer. New knowledge the night before unsettles the pairs that had already set.
Traps
First, the back of the card. Do not paste an AI definition without checking it. Two of our ten needed correcting, both in law, and both on the difference between a rule and its exception or between publishing and withholding. A wrong card recalled at spaced intervals installs the error more firmly than not carding it at all.
Second, the tag. Carrying the major item over from the log is fast and inherits any error in the log. A mistagged card breaks the filtered review by major item, which is the main reason to tag at all.
Third, volume. It is tempting to card everything, and the review load then exceeds the ten minutes a day the final fortnight can spare. The point of running only your mistakes is the small number. Cap the new cards.
Fourth, the front. Because Anki matches on the first field, rewording the front breaks the link to the existing card and its history. Fix the wording of the front when you first create the card and leave it alone.
Summary
In the final fortnight, run your mistakes rather than everything. Pour every stage into one four-column table, have the AI aggregate by major item and error type and draft the cards, check the backs against primary sources, and import through a CSV with header directives. Because Anki updates on a matching first field by default, you can rebuild and re-import weekly without duplicating cards or losing the schedule. Ten minutes a day of review, thirty minutes at the weekend to refresh, and no new cards the night before.
That closes the six methods. What remains is the domain where the material itself goes stale: bringing law and ethics up to the current state of the rules before you memorise anything.
Sources
- Anki manual, Importing Text Files (duplicate handling and preserved scheduling)
- Anki manual, Deck Options (New Cards/Day)
- Personal Information Protection Commission, anonymised information
- Personal Information Protection Commission, FAQ on pseudonymised information
- JDLA, About the G-test (where to get Syllabus 2024)





