05Agentic crawl & retrievalClient work
Admissions RAG
An AI crawler that reads Italian universities' admission rules, and answers only when it can quote the evidence.
- Role
- Sole developer, every commit
- Built
- January to April 2026
- Type
- Client work
- Rule
- Evidence, or no answer
In short
Italian universities publish their admission rules inconsistently. This system crawls each university until an AI judge agrees the admission calls have been found, extracts every programme's rules in one long read that must quote its evidence or return nothing, and feeds a chatbot that cites the exact pages it used.
The same admission rule hides under different names on every university's website.
One calls it a bando, another an immatricolazione procedure, a third hides it in a student guide PDF, and several use bando mostly for PhD calls or prizes. Keyword scrapers either miss the real call or drown in irrelevant PDFs, and a chatbot built on guesses would tell students the wrong thing about entrance tests.
What it does
Searches until it's sure
Discovery runs in rounds: an AI coverage judge asks whether an undergraduate admission call is in the list, and for gaps new queries are proposed and landing pages read for PDF links.
Judges PDFs cheaply
Free heuristics on the address, file name and link text settle most PDFs; a fast model looks only at the ambiguous middle.
Evidence or null
Each programme's admission type and teaching language come with a confidence score, a direct quote and a source link. Below 0.6 confidence the answer is null.
Watches for changes
Tracked pages are re-checked, and diffs are analysed for changes that matter to admissions.
Answers with citations
The chatbot rewrites follow-up questions, retrieves with a university filter, answers from numbered sources, and a guard checks every citation.
Sole engineer
Both stages, from the discovery loop and the extraction rules to the chat API, its retrieval and its citation guard.
Stage one turns messy university websites into structured, evidenced rules. Stage two answers questions about them and shows its sources.
Find, extract, answer
Both stages, and the store that connects them.
What each part does
- One long-context read per programme
- One long-context Gemini call per programme, with all its pages at once.
- Chat API
- The chat endpoint.
- Discovery: search and site maps
- Searches and maps each university's site for admission pages and PDFs.
- Embeddings
- Turns the corpus into vectors.
- Write the answer
- Gemini writes the answer from numbered sources.
- Citation check
- Checks every citation marker points at a real retrieved passage.
- AI coverage judge
- An AI check: is there an undergraduate admission call in what we found?
- Rules and pages
- Structured rules with quotes and sources, plus the page corpus.
- Landing-page reader
- Reads admission landing pages and pulls out the PDF links on them.
- Null, never a guess
- No evidence means no answer, never a guess.
- Two-tier PDF judge
- Heuristics first, a fast model only for the unclear cases.
- Postgres with pgvector
- Postgres with pgvector.
- New search queries
- For gaps, Gemini proposes fresh search queries.
- Retrieve, filtered by university
- Finds the closest passages, filtered to one university when asked.
- Rewrite the question
- Turns a follow-up question into one that stands on its own.
- Scrape the pages
- Fetches the chosen pages and documents as text.
- Programme list
- The list of programmes to cover.
- Student or advisor
- A student or an admissions advisor.
- Change watch
- Re-checks tracked pages and flags admission-relevant changes.
Discovery, round by round
How the crawler keeps searching one university until the judge is satisfied.
1 of 15
Round one starts for a university.
Discovery loop → Discovery: Round 1 for this university
All 15 steps as text
What each part does
- Discovery loop
- The loop that runs the rounds.
- Discovery
- The discovery service.
- Firecrawl
- Firecrawl search and site mapping.
- Coverage judge
- The AI coverage judge.
- Query writer
- Proposes new search queries.
- Landing reader
- Reads admission landing pages for PDF links.
A question, with audited citations
How the chatbot answers, and how it checks itself.
1 of 10
A question arrives, with the conversation and optionally a university.
User → Chat API: Question, history, optional university
All 10 steps as text
What each part does
- User
- A student or an advisor.
- Chat API
- The chat API.
- Query rewriter
- The query rewriter.
- Retrieval
- Vector retrieval in pgvector.
- Gemini
- Gemini, writing the answer.
- Citation check
- The citation guard.
Decisions and trade-offs
Scrape deeply, analyse once
A rewrite replaced a first version full of regex patterns and fallbacks with one long-context Gemini call per programme, reading all of its pages at once.
Null over guess
Evidence is mandatory and low confidence returns null. The confidence scale is written down: an explicit statement, a strong inference or a moderate one.
An AI check instead of keyword rules
Universities name their calls too inconsistently for keywords, so coverage is judged by a model. One flaky call is logged and skipped, never allowed to stop a multi-university run.
Cheap first, clever second
Heuristics settle most PDFs before any model is asked, and change tracking runs in a mode that costs no extra crawling.
Citations audited, not enforced
The guard reports uncited sentences as warnings instead of rejecting the answer, because a short uncited summary line is acceptable to readers.
Results
- 0.6confidence floorbelow it, the answer is null
- 2stagesan agentic crawl, then a cited chatbot
- 104backend test functionsin 12 test files
- 41commits, all hisJanuary to April 2026
Stack
- Crawl and extraction
- Python
- Firecrawl
- Gemini
- Retrieval
- Gemini embeddings
- Postgres with pgvector
- Service
- FastAPI
- Pydantic v2
- Alembic
- pytest
- Docker