Skip to content
fethiverseHire me

03Agentic RAGClient work

The Faculty

An AI pipeline that writes exam questions, then tries to break every one before it ships.

Role
Sole developer, every commit
Built
February to July 2026
Type
Client work
Output
Validated multiple-choice questions

In short

Built for an Italian exam-preparation app, this pipeline turns source PDFs into multiple-choice questions. Specialised agents read the material, draft a question and engineer plausible wrong answers, then attack the result: a blind solver, checks against the source, and a second AI model whenever the first is unsure. Only questions that survive, or are precisely repaired, ship.

Exam-prep apps need large numbers of new questions, and a plain AI model writes bad ones.

Straight generation produces questions that are ambiguous, have two defensible answers, give the answer away through option length, or state facts that are not in the source. The questions have to match a real exam's style, difficulty and syllabus, with wrong answers that are wrong for the right reasons.

What it does

  • Reads the material

    Parses source PDFs and past questions, extracts concepts and their hierarchy, and aligns them with the syllabus.

  • Plans every batch

    Chooses topic, Bloom level and question type per slot, with a ledger that stops the same idea appearing twice in a batch.

  • Writes the wrong answers on purpose

    After a self-critique, a distractor agent writes options that reflect the mistakes real students make.

  • Attacks every question

    A solver answers blind with shuffled options, grounding and fact checks trace each claim to the source, and 16 rule-based validators check balance and structure.

  • Repairs instead of restarting

    A failed question gets a targeted repair, then a semantic duplicate check against everything written before.

Sole engineer

Architecture, agents, validators, the API and its container: all of the repository's commits are Fethi's.

Generation is the easy part. Most of the system exists to prove each question is right before anyone sees it.

From a PDF to a question you can trust

The pipeline end to end, with the attack stage in the middle.

Try to break itpassesfailsVerdictSelf-critiqueSecond model, when unsureDuplicate check16 rule-based validatorsEngineer the wrong answersDraft the questionGrounding and fact checksValidated question withexplanationRead: concepts and syllabusPlan the batch: topic, level,typeSurgical repairBlind solver, shuffled optionsSource PDFs and pastquestions
What each part does
Verdict
Combines every check into one verdict.
Self-critique
The agent reviews its own draft before going further.
Second model, when unsure
A second Gemini model re-solves only the uncertain questions, which keeps cost down. Disagreement flags ambiguity.
Duplicate check
A semantic similarity check against earlier questions.
16 rule-based validators
Sixteen deterministic checks: option length balance, structure, blocklists, metadata, language.
Engineer the wrong answers
Writes wrong answers that are plausible for the right reasons.
Draft the question
Drafts the question and its correct answer.
Grounding and fact checks
Checks that every claim can be traced back to the source material.
Validated question with explanation
The question ships with its explanation and metadata.
Read: concepts and syllabus
PDF parsing, concept extraction and alignment with the syllabus.
Plan the batch: topic, level, type
Decides what each question slot should test, and prevents repeats.
Surgical repair
A targeted fix of what failed, instead of starting over.
Blind solver, shuffled options
Answers the question without knowing the key, with the options shuffled. If it cannot find the answer, neither can a student.
Source PDFs and past questions
The course material and existing question banks the client provides.

One question, attacked before it ships

Follow a single question slot through writing, attack and repair.

opt · The critique failsopt · Confidence is lowalt · The verdict failsIt passesSlot: topic, level, syllabus context1Question and correct answer2Critique the draft3Issues, or a pass4Rewrite with the critique5Engineer the wrong answers6Plausible distractors7Validate8Solve it blind, options shuffled9Chosen answer and confidence10Solve it again with another model11Agree or disagree12Combined verdict13Targeted repair instruction14Repaired question15Validate again16Duplicate check, then save17WorkerQuestion writerSelf-critiqueDistractor engineerValidatorBlind solverSecond modelRepair
All 17 steps as text
What each part does
Worker
The worker that runs one question slot from start to finish.
Question writer
The question-writing agent.
Self-critique
The self-critique agent.
Distractor engineer
The distractor engineer: writes the wrong answers.
Validator
The composite validator that runs every check.
Blind solver
The blind solver.
Second model
A second Gemini model, used only when the first is unsure.
Repair
The repair agent.

Decisions and trade-offs

  1. Attack the output instead of trusting the generator

    Blind solving with shuffled options catches answers that give themselves away; grounding checks catch facts that are not in the source.

  2. A second model only when it matters

    Cross-model validation runs only on low-confidence questions, which keeps cost down while still catching ambiguity.

  3. Wrap, don't rewrite

    The migration to Google ADK, in progress, wraps the existing stages as agents in deterministic workflow graphs. The validators and repair logic are kept as they are, and state passes through session state instead of shared mutation.

  4. Validators over temperature

    The current Gemini models expect temperature 1.0, so reliability comes from the checks around the model rather than from turning its randomness down.

  5. One door to the model

    Every model call goes through a single client with a rate limiter and a circuit breaker, so a bad minute at the provider cannot cascade.

Results

  • 2,461test functionsin 155 test files
  • 16rule-based validatorsplus the agent checks
  • 236commits, all hisFebruary to July 2026
  • ~20agent modules

Stack

AI
  • Gemini
  • Google ADK (in migration)
  • Multi-agent validation
Retrieval
  • ChromaDB
  • sentence-transformers
  • PyMuPDF
Service
  • Python
  • FastAPI
  • Docker
  • pytest