知識がなくても始められる、AIと共にある豊かな毎日。
AI Learning and Certification

Let AI Write the Questions, Then Check Them: Measuring the Error Rate Across 80 Items

swiftwand

There are only twenty official sample questions. Buy one problem set and the questions covering the major item you are weakest on run out after a handful. So you ask an AI to write G-test questions, and plausible four-option items appear without limit. But is the marked answer actually right? Does the explanation agree with the primary source? Solve them without knowing, memorise a wrong answer, and the time you spent solving works against you.

This article contains affiliate links. As an Amazon Associate, we earn from qualifying purchases.

This article answers that worry with numbers. This site had Claude write questions for all ten major items of the G-test, 80 in total, with and without the syllabus, and checked every one against primary material: the official samples, the syllabus, statutes and guidelines. The headline: of the 50 written with the syllabus, one needed a correction; of the 30 written without it, twelve did. The errors clustered on proper nouns in law and guidelines. What follows is where the gap comes from, what to do with the questions that pass, and how to build verification into the routine.

忍者AdMax

Decide three things before you generate

First, the format. The twenty official samples break down into four shapes, all four-option: choose the single most appropriate option (12, including the expected-value calculation), choose the single least appropriate option (4), fill the blank (3), and choose the combination of terms (1). The least-appropriate shape is used in all four law and ethics samples. Write those four shapes into the instruction and the generated questions line up with the real thing. In our measurement we forgot to include the least-appropriate shape, so every generated law and ethics question came out as most-appropriate. That is a limitation of the measurement, noted below.

Second, the unit. Ask for G-test questions as a lump and the AI drifts towards what it knows best. Ask per major item, and hand over the mid-level items and keywords from the syllabus, and the scope is fixed. The knowledge base built in the previous step earns its keep here. As covered in the article on ingesting the syllabus, the AI does not know the exam scope by itself.

Third, the shape of the output. Require the stem, the options, the answer, a one-sentence explanation, the central term and the name of the mid-level item, under fixed field names. That lets you check mechanically, later, whether the term exists in the syllabus and whether the mid-level item name is real. We asked for JSON; a table works too.

The prompt, with the syllabus attached

You are an assistant helping someone study for the JDLA G-test.
Write 5 practice questions for the major item Overview of machine learning.
Match the shapes used in the official JDLA samples: four options A to D,
choose the single most appropriate, choose the single least appropriate,
fill the blank, and choose the combination of terms.
Restrict the scope strictly to the syllabus below (objectives and keywords).
Do not make a term that is not in the keywords the correct answer.
----- syllabus -----
(paste the objectives and keywords for that major item)
----- end -----
For each question give the stem, options A to D, the answer, a one-sentence
explanation, the central term, and the name of the mid-level item.

The sentence about not making an out-of-keyword term the answer is what makes the scope bite. Without the syllabus attached, writing that sentence is empty, because the AI has nothing to check against.

Attaching the syllabus is a prompt design choice. For a practical guide to structuring prompts and context so a model stays inside your source material, this book covers it in depth.

USD 54.51 on Amazon.com (as of 2026/09/22)

How the measurement was built: 80 questions on four axes

Two conditions. Under A the syllabus went in and the model wrote five questions per major item, 50 in total. Under B only the major item name and the format went in, three per item, 30 in total. The model was claude-sonnet-5 with no tools and no project configuration, to approximate a learner using Claude as it comes. Other vendors were not compared, as this site has no API contract with them.

AxisMethodError types
ScopeDoes the central term exist in the keyword list of revision 1.4 (matched with parentheticals and the suffix meaning definition removed)? Does it avoid terms deleted between the original Syllabus 2024 of May 2024 and revision 1.4, such as ENIAC, fifth-generation computers and ChatGPT? The revision history lists examples, not an exhaustive set.Term not in the keyword list; use of a deleted term
StructureDoes the named mid-level item exist, and is the major item assignment correct?Wrong major item
AnswerDoes the marked answer contradict a primary source: the syllabus objective, the official samples, or the text of a statute or guideline? Is more than one option defensible?Wrong answer; wrong fact or attribution
WordingDo the stem, the correct option and the explanation contain self-contradiction or inaccuracy?Minor inaccuracy

Scope and structure can be judged mechanically against the keyword list. Answer and wording were judged by reading all 80. For the legal items we referred to the Act on the Protection of Personal Information, article 30-4 of the Copyright Act, the contract guideline on the use of AI and data from the Ministry of Economy, Trade and Industry, and the AI utilisation guideline from the Ministry of Internal Affairs and Communications.

Result: 2 percent against 40 percent

ConditionQuestionsWrong answer or factAny correction neededTerm in keyword listMid-level item name realWrong major item
A: with syllabus501 (2.0%)1 (2.0%)50 of 5045 of 500
B: without syllabus302 (6.7%)12 (40%)24 of 303 of 304

The 45 under condition A reflects five statistics questions that returned the text of an objective, such as can calculate basic statistics, instead of the mid-level item name; the major item assignment was correct. The 24 under condition B is the figure after removing parentheticals and the definition suffix; of the remaining six, one (article 30-4 of the Copyright Act) fell out of the mechanical match on a notation difference, and five took a term absent from the keyword list as their subject.

All 50 questions under condition A used terms present in the syllabus keywords and were assigned to the right major item. One needed a correction, where the wording of the correct option diverged from the sense of the primary source and contradicted itself. Under condition B only 3 of 30 mid-level item names matched a real one: the AI was classifying under a taxonomy of its own invention.

Splitting the twelve flagged questions by type:

Error typeCountExample
Wrong answer1The name of the contract method diverged from the term used in the primary source
Wrong fact or attribution1Assigned a guideline to the wrong ministry
Minor inaccuracy2Vague attribution of who proposed a concept; oversimplified historical causation
Term not in the keyword list as the subject4Narrow and general artificial intelligence, the first AI boom, Bayes theorem, the OECD AI Principles
Wrong major item4Dropout and learning rate filed under component techniques; contracts and bias filed under social implementation
Use of a deleted term1ENIAC among the options

The counts sum to 13, one more than 12, because the contract-guideline question is both a wrong answer and a wrong major item. Note that not in the keyword list is not the same as out of the exam scope. General artificial intelligence and the AI booms appear in the text of the syllabus objectives, and the official samples include a question on the first and second booms. Being absent from the keyword list means the term falls outside your glossary and your weakness-tracking axis; it does not mean it will not be examined.

By major item, the flags under condition B fell three on what artificial intelligence is, three on social implementation, two on component techniques, and one each on trends, statistics, ethics and governance, and the overview of deep learning. The three on what artificial intelligence is came from keyword-absent terms and vague attribution; on social implementation, two were a contract question that belongs in the legal domain and a bias question that belongs in ethics, with the third being a guideline filed under the wrong ministry. The AI writes questions by association from the name of the major item, so the broader the name, the more the scope drifts.

Worth noticing: the technical questions were almost entirely correct under both conditions. Questions on convolution layers, batch normalisation, skip connections, word2vec and CycleGAN had answers and explanations that did not contradict any primary source. The errors were in proper nouns from law, guidelines and history.

A gap of 2 percent against 40 percent only shows up because the questions were checked one by one. For a structured approach to evaluating model output, this book on building with foundation models is a good reference.

Three of the errors, in outline

The questions themselves are not reproduced; only the substance.

The first came from condition B under social implementation, and asked for the name of the contract method recommended by the contract guideline on the use of AI and data from the Ministry of Economy, Trade and Industry. The AI marked multi-stage contract as correct and put exploratory staged contract among the distractors. The method the guideline proposes is named the exploratory staged method, in which the contract is split across assessment, proof of concept, development and additional training and concluded stage by stage. There is no method called multi-stage contract in the guideline. In other words, the option the AI marked wrong is the one bearing the name used in the primary source; the answer and the distractor had swapped places. The explanation was fluent, and you would not notice without knowing.

The second, also from condition B, asked in a fill-the-blank for the name of a guideline formulated by the Ministry of Economy, Trade and Industry, and marked the AI Utilisation Guideline as correct. That guideline was published in August 2019 by the AI Network Society Promotion Council of the Ministry of Internal Affairs and Communications, not by METI. And the current consolidated document is now the AI Business Operator Guidelines, revision 1.2 dated 31 March 2026, published jointly by both ministries. A wrong ministry and a stale revision in one question.

The third is the single flagged question under condition A, on the thinking in the Camera Image Utilisation Guidebook. The correct option said, in substance, that after acquiring images you publish or post the purpose of use and the installation situation as necessary, and then went on to say that privacy consideration should be built in from the planning and design stage. The first half and the second half contradict each other. The guidebook is built on advance notice before filming and on consideration from the design stage, so the first half diverges from the primary source. The direction of the option is right, but because the sentence marked as correct does not match the source, we counted it as a wrong answer. Handing over the syllabus does not guarantee the quality of the prose.

Why law and proper nouns

What the three share is that correctness turns on proper nouns, on which body owns a document and on which revision is current, rather than on understanding a concept. Technical terms have their definitions restated in text all over the world, so what the AI knows and what the exam wants tend to coincide. The name of a guideline issued by a Japanese ministry, or the name of a method inside it, appears in far less text, sits alongside documents with similar names, and gets consolidated or renamed at revision. The AI picks the most plausible-sounding name and marks it correct.

The reason handing over the syllabus reduces this has the same structure. The keyword list contains the exact strings, the contract guideline on the use of AI and data, the Camera Image Utilisation Guidebook, so the model has something to copy from rather than something to reconstruct.

What to do with the questions that pass

Do not solve them on the spot. Leave them a day and then solve under the 41-second limit. Right after generating you still remember the answers, so solving proves nothing.

The order is the twenty official samples first, generated questions second. Learn the shape from the samples, then use generated questions for volume. Generated questions are not the same as the real thing. They are practice built within the syllabus scope, with no guarantee that the difficulty or the construction of the distractors matches the exam. Use them knowing that.

One example that passed verification, in outline: a fill-the-blank from condition A under what artificial intelligence is, asking for the phenomenon whereby a newly realised AI technique comes to be seen as mere automation and is dropped from what counts as AI. The answer was the AI effect, with singularity, the frame problem and symbol grounding among the options. Every term is a keyword from mid-level items 1 and 2, and the distractors come from the same items. A question like that can be used as written.

This article does not include a question bank, for two reasons. G-test questions are not published as a rule, and the only official past questions are the JDLA twenty. Lining up lookalikes and calling them past papers runs against the reason they are not published. The other reason is that the value lies in generating against your own weak major item. Five questions on the mid-level item you are weak on beat thirty written by someone else.

Cost and time

The twenty calls in the measurement took about five and a half minutes and cost 0.35 dollars. Five questions for one major item is one call, a dozen seconds, a few yen. Generation is free in practical terms; verification is what costs.

Verification runs tens of times longer than generation. But if you clear scope and structure mechanically first, the reading time concentrates on the handful of legal questions. Across our 80, the technical questions were settled by matching against the keyword list, and the reading was needed on law and history. For a 45-minute day, five questions for one major item in one call plus verification is 30 minutes, and you solve them the next day.

Traps

First, how you hand over the samples. We forgot to include the least-appropriate shape in the instruction, so all the generated law and ethics questions came out as most-appropriate, while every one of the four official law and ethics samples uses the least-appropriate shape. If the four shapes are not in the instruction, the generated set will not match the exam.

Second, coverage. The AI will produce as many questions as you ask for on the major items it finds easy. Generating per major item, in equal numbers, is what keeps the coverage even. Spending 40 days only on the items you enjoy is a way of manufacturing your own blind spot.

Third, treating a passed question as certified. Verification confirms that nothing contradicts the primary sources we checked. It does not confirm that the question is a good question, or that it is at the level of the real exam.

Summary

Generated questions are usable, on the condition that you hand over the syllabus and verify. With the syllabus, one of 50 needed a correction; without it, twelve of 30, and only three of 30 named a real mid-level item. The errors are concentrated in law, guidelines and history, where correctness depends on proper nouns, ownership and revisions. Clear scope and structure mechanically, read the legal ones yourself, and solve the survivors a day later at 41 seconds each.

Sources

ブラウザだけでできる本格的なAI画像生成【ConoHa AI Canvas】
ABOUT ME
swiftwand
swiftwand
AIを使って、毎日の生活をもっと快適にするアイデアや将来像を発信しています。 初心者にもわかりやすく、すぐに取り入れられる実践的な情報をお届けします。 Sharing ideas and visions for a better daily life with AI. Practical tips that anyone can start using right away.
記事URLをコピーしました