Skip to content
fethiverseHire me

05Agentic crawl & retrievalClient work

Admissions RAG

An AI crawler that reads Italian universities' admission rules, and answers only when it can quote the evidence.

Role
Sole developer, every commit
Built
January to April 2026
Type
Client work
Rule
Evidence, or no answer

In short

Italian universities publish their admission rules inconsistently. This system crawls each university until an AI judge agrees the admission calls have been found, extracts every programme's rules in one long read that must quote its evidence or return nothing, and feeds a chatbot that cites the exact pages it used.

The same admission rule hides under different names on every university's website.

One calls it a bando, another an immatricolazione procedure, a third hides it in a student guide PDF, and several use bando mostly for PhD calls or prizes. Keyword scrapers either miss the real call or drown in irrelevant PDFs, and a chatbot built on guesses would tell students the wrong thing about entrance tests.

What it does

  • Searches until it's sure

    Discovery runs in rounds: an AI coverage judge asks whether an undergraduate admission call is in the list, and for gaps new queries are proposed and landing pages read for PDF links.

  • Judges PDFs cheaply

    Free heuristics on the address, file name and link text settle most PDFs; a fast model looks only at the ambiguous middle.

  • Evidence or null

    Each programme's admission type and teaching language come with a confidence score, a direct quote and a source link. Below 0.6 confidence the answer is null.

  • Watches for changes

    Tracked pages are re-checked, and diffs are analysed for changes that matter to admissions.

  • Answers with citations

    The chatbot rewrites follow-up questions, retrieves with a university filter, answers from numbered sources, and a guard checks every citation.

Sole engineer

Both stages, from the discovery loop and the extraction rules to the chat API, its retrieval and its citation guard.

Stage one turns messy university websites into structured, evidenced rules. Stage two answers questions about them and shows its sources.

Find, extract, answer

Both stages, and the store that connects them.

Stage 1: find and extractStage 2: answerquoted evidenceno evidencegapcoveredgapOne long-context read perprogrammeChat APIDiscovery: search and sitemapsEmbeddingsWrite the answerCitation checkAI coverage judgeRules and pagesLanding-page readerNull, never a guessTwo-tier PDF judgePostgres with pgvectorNew search queriesRetrieve, filtered by universityRewrite the questionScrape the pagesProgramme listStudent or advisorChange watch
What each part does
One long-context read per programme
One long-context Gemini call per programme, with all its pages at once.
Chat API
The chat endpoint.
Discovery: search and site maps
Searches and maps each university's site for admission pages and PDFs.
Embeddings
Turns the corpus into vectors.
Write the answer
Gemini writes the answer from numbered sources.
Citation check
Checks every citation marker points at a real retrieved passage.
AI coverage judge
An AI check: is there an undergraduate admission call in what we found?
Rules and pages
Structured rules with quotes and sources, plus the page corpus.
Landing-page reader
Reads admission landing pages and pulls out the PDF links on them.
Null, never a guess
No evidence means no answer, never a guess.
Two-tier PDF judge
Heuristics first, a fast model only for the unclear cases.
Postgres with pgvector
Postgres with pgvector.
New search queries
For gaps, Gemini proposes fresh search queries.
Retrieve, filtered by university
Finds the closest passages, filtered to one university when asked.
Rewrite the question
Turns a follow-up question into one that stands on its own.
Scrape the pages
Fetches the chosen pages and documents as text.
Programme list
The list of programmes to cover.
Student or advisor
A student or an admissions advisor.
Change watch
Re-checks tracked pages and flags admission-relevant changes.

Discovery, round by round

How the crawler keeps searching one university until the judge is satisfied.

loop · Until covered, or out of roundsalt · CoveredA gapparRound 1 for this university1Search and map the site2Candidate pages and PDFs3Discovery result4Is an undergraduate admission call inthis list?5Verdict6Stop7Propose new queries8Fresh queries9Search again10More candidates11Read the admission landing pages12PDF links found on them13Merge the candidates14Final result, flagged if still missing15Discovery loopDiscoveryFirecrawlCoverage judgeQuery writerLanding reader
All 15 steps as text
What each part does
Discovery loop
The loop that runs the rounds.
Discovery
The discovery service.
Firecrawl
Firecrawl search and site mapping.
Coverage judge
The AI coverage judge.
Query writer
Proposes new search queries.
Landing reader
Reads admission landing pages for PDF links.

A question, with audited citations

How the chatbot answers, and how it checks itself.

opt · Rewriting is onQuestion, history, optional university1History and question2A standalone question, or the originalon failure3Embed and search, with filters4Passages with page links5Answer from these numbered sources6Answer with source markers7Check every marker against thepassages8Warnings, if any9Answer, citations and warnings10UserChat APIQuery rewriterRetrievalGeminiCitation check
All 10 steps as text
What each part does
User
A student or an advisor.
Chat API
The chat API.
Query rewriter
The query rewriter.
Retrieval
Vector retrieval in pgvector.
Gemini
Gemini, writing the answer.
Citation check
The citation guard.

Decisions and trade-offs

  1. Scrape deeply, analyse once

    A rewrite replaced a first version full of regex patterns and fallbacks with one long-context Gemini call per programme, reading all of its pages at once.

  2. Null over guess

    Evidence is mandatory and low confidence returns null. The confidence scale is written down: an explicit statement, a strong inference or a moderate one.

  3. An AI check instead of keyword rules

    Universities name their calls too inconsistently for keywords, so coverage is judged by a model. One flaky call is logged and skipped, never allowed to stop a multi-university run.

  4. Cheap first, clever second

    Heuristics settle most PDFs before any model is asked, and change tracking runs in a mode that costs no extra crawling.

  5. Citations audited, not enforced

    The guard reports uncited sentences as warnings instead of rejecting the answer, because a short uncited summary line is acceptable to readers.

Results

  • 0.6confidence floorbelow it, the answer is null
  • 2stagesan agentic crawl, then a cited chatbot
  • 104backend test functionsin 12 test files
  • 41commits, all hisJanuary to April 2026

Stack

Crawl and extraction
  • Python
  • Firecrawl
  • Gemini
Retrieval
  • Gemini embeddings
  • Postgres with pgvector
Service
  • FastAPI
  • Pydantic v2
  • Alembic
  • pytest
  • Docker