03Agentic RAGClient work
The Faculty
An AI pipeline that writes exam questions, then tries to break every one before it ships.
- Role
- Sole developer, every commit
- Built
- February to July 2026
- Type
- Client work
- Output
- Validated multiple-choice questions
In short
Built for an Italian exam-preparation app, this pipeline turns source PDFs into multiple-choice questions. Specialised agents read the material, draft a question and engineer plausible wrong answers, then attack the result: a blind solver, checks against the source, and a second AI model whenever the first is unsure. Only questions that survive, or are precisely repaired, ship.
Exam-prep apps need large numbers of new questions, and a plain AI model writes bad ones.
Straight generation produces questions that are ambiguous, have two defensible answers, give the answer away through option length, or state facts that are not in the source. The questions have to match a real exam's style, difficulty and syllabus, with wrong answers that are wrong for the right reasons.
What it does
Reads the material
Parses source PDFs and past questions, extracts concepts and their hierarchy, and aligns them with the syllabus.
Plans every batch
Chooses topic, Bloom level and question type per slot, with a ledger that stops the same idea appearing twice in a batch.
Writes the wrong answers on purpose
After a self-critique, a distractor agent writes options that reflect the mistakes real students make.
Attacks every question
A solver answers blind with shuffled options, grounding and fact checks trace each claim to the source, and 16 rule-based validators check balance and structure.
Repairs instead of restarting
A failed question gets a targeted repair, then a semantic duplicate check against everything written before.
Sole engineer
Architecture, agents, validators, the API and its container: all of the repository's commits are Fethi's.
Generation is the easy part. Most of the system exists to prove each question is right before anyone sees it.
From a PDF to a question you can trust
The pipeline end to end, with the attack stage in the middle.
What each part does
- Verdict
- Combines every check into one verdict.
- Self-critique
- The agent reviews its own draft before going further.
- Second model, when unsure
- A second Gemini model re-solves only the uncertain questions, which keeps cost down. Disagreement flags ambiguity.
- Duplicate check
- A semantic similarity check against earlier questions.
- 16 rule-based validators
- Sixteen deterministic checks: option length balance, structure, blocklists, metadata, language.
- Engineer the wrong answers
- Writes wrong answers that are plausible for the right reasons.
- Draft the question
- Drafts the question and its correct answer.
- Grounding and fact checks
- Checks that every claim can be traced back to the source material.
- Validated question with explanation
- The question ships with its explanation and metadata.
- Read: concepts and syllabus
- PDF parsing, concept extraction and alignment with the syllabus.
- Plan the batch: topic, level, type
- Decides what each question slot should test, and prevents repeats.
- Surgical repair
- A targeted fix of what failed, instead of starting over.
- Blind solver, shuffled options
- Answers the question without knowing the key, with the options shuffled. If it cannot find the answer, neither can a student.
- Source PDFs and past questions
- The course material and existing question banks the client provides.
One question, attacked before it ships
Follow a single question slot through writing, attack and repair.
1 of 17
A slot arrives: what to test, at what level, from which part of the syllabus.
Worker → Question writer: Slot: topic, level, syllabus context
All 17 steps as text
What each part does
- Worker
- The worker that runs one question slot from start to finish.
- Question writer
- The question-writing agent.
- Self-critique
- The self-critique agent.
- Distractor engineer
- The distractor engineer: writes the wrong answers.
- Validator
- The composite validator that runs every check.
- Blind solver
- The blind solver.
- Second model
- A second Gemini model, used only when the first is unsure.
- Repair
- The repair agent.
Decisions and trade-offs
Attack the output instead of trusting the generator
Blind solving with shuffled options catches answers that give themselves away; grounding checks catch facts that are not in the source.
A second model only when it matters
Cross-model validation runs only on low-confidence questions, which keeps cost down while still catching ambiguity.
Wrap, don't rewrite
The migration to Google ADK, in progress, wraps the existing stages as agents in deterministic workflow graphs. The validators and repair logic are kept as they are, and state passes through session state instead of shared mutation.
Validators over temperature
The current Gemini models expect temperature 1.0, so reliability comes from the checks around the model rather than from turning its randomness down.
One door to the model
Every model call goes through a single client with a rate limiter and a circuit breaker, so a bad minute at the provider cannot cascade.
Results
- 2,461test functionsin 155 test files
- 16rule-based validatorsplus the agent checks
- 236commits, all hisFebruary to July 2026
- ~20agent modules
Stack
- AI
- Gemini
- Google ADK (in migration)
- Multi-agent validation
- Retrieval
- ChromaDB
- sentence-transformers
- PyMuPDF
- Service
- Python
- FastAPI
- Docker
- pytest