06Benchmark comparison app
LLMetric
An app that puts two AI models side by side on one chart, so you can see which is better at what.
- Role
- Design and build
- Built
- Mostly April to May 2025
- Status
- Prototype, demo data
- Stack
- Next.js and Supabase
In short
LLMetric compares large language models head to head: a model catalogue, versus pages with a five-axis radar chart, and a news section with its own editor. It is Fethi's earliest project here. The model data is demo data, and keeping it current was explored with an AgentQL extraction prototype.
Every AI model launch quotes different benchmarks in a different format.
Comparing two models means reading several pages and converting numbers in your head. LLMetric puts scores on one comparable scale and shows two models on a single chart.
What it does
A versus view
Pick two models; both appear on one pentagon radar chart across reasoning, knowledge, math, coding and creativity, with a list of what each does better.
A model catalogue
Browse models and open a detail page for each.
News with an editor
An articles section with an admin area to write and edit posts, stored in Supabase.
An extraction prototype
An AgentQL worker that pulls structured data out of a web page from a natural-language query, as a first step toward live benchmark data.
An early build
His first product of this kind: he designed and built the interface, the charts, the data model and the extraction prototype.
A Next.js app on Supabase. The model data is static demo data today; the extraction worker shows how it could be kept current.
The app as it stands
What is live in the code today, and what is a prototype.
What each part does
- Article editor
- An admin area to create and edit articles.
- AgentQL
- AgentQL's extraction service.
- Demo model data
- Static demo model data, not scraped.
- Editor
- Whoever writes the news articles.
- AgentQL extraction worker
- A proof of concept: extract structured data from a page with a natural-language query.
- Models, compare, articles
- The model catalogue, the compare page and the articles.
- Radar chart, five axes
- A Recharts radar chart with five axes.
- Supabase: articles
- Supabase stores the articles.
- Visitor
- Anyone comparing models.
- Next.js app
- A Next.js App Router app with shadcn/ui components.
Comparing two models
The versus view, step by step.
1 of 6
The visitor picks two models.
Visitor → Compare page: Pick two models
All 6 steps as text
What each part does
- Visitor
- A visitor.
- Compare page
- The compare page.
- Model data
- The model data.
- Radar chart
- The radar chart.
Decisions and trade-offs
From a custom backend to a managed one
It started as a monorepo with a FastAPI backend, ETL workers, Postgres and Redis in Docker, and was simplified to a Next.js app on Supabase: less to run, at the cost of the custom API.
A versus view over tables
A radar chart with two overlaid models was chosen over tabbed tables after several interface iterations, because the difference is visible at a glance.
Say what is real
The model data is demo data and the extraction worker is a prototype. This page says so, rather than implying a live pipeline.
Results
- 5comparison axesreasoning, knowledge, math, coding, creativity
- 26commits25 of them in two weeks of 2025
- 1stproject of this setwhere the later work started
Stack
- App
- Next.js
- TypeScript
- Tailwind CSS
- shadcn/ui
- Recharts
- Data
- Supabase
- AgentQL (prototype)