Skip to content
fethiverseHire me

06Benchmark comparison app

LLMetric

An app that puts two AI models side by side on one chart, so you can see which is better at what.

Role
Design and build
Built
Mostly April to May 2025
Status
Prototype, demo data
Stack
Next.js and Supabase

In short

LLMetric compares large language models head to head: a model catalogue, versus pages with a five-axis radar chart, and a news section with its own editor. It is Fethi's earliest project here. The model data is demo data, and keeping it current was explored with an AgentQL extraction prototype.

Every AI model launch quotes different benchmarks in a different format.

Comparing two models means reading several pages and converting numbers in your head. LLMetric puts scores on one comparable scale and shows two models on a single chart.

What it does

  • A versus view

    Pick two models; both appear on one pentagon radar chart across reasoning, knowledge, math, coding and creativity, with a list of what each does better.

  • A model catalogue

    Browse models and open a detail page for each.

  • News with an editor

    An articles section with an admin area to write and edit posts, stored in Supabase.

  • An extraction prototype

    An AgentQL worker that pulls structured data out of a web page from a natural-language query, as a first step toward live benchmark data.

An early build

His first product of this kind: he designed and built the interface, the charts, the data model and the extraction prototype.

A Next.js app on Supabase. The model data is static demo data today; the extraction worker shows how it could be kept current.

The app as it stands

What is live in the code today, and what is a prototype.

Prototypenatural-language queryArticle editorAgentQLDemo model dataEditorAgentQL extraction workerModels, compare, articlesRadar chart, five axesSupabase: articlesVisitorNext.js app
What each part does
Article editor
An admin area to create and edit articles.
AgentQL
AgentQL's extraction service.
Demo model data
Static demo model data, not scraped.
Editor
Whoever writes the news articles.
AgentQL extraction worker
A proof of concept: extract structured data from a page with a natural-language query.
Models, compare, articles
The model catalogue, the compare page and the articles.
Radar chart, five axes
A Recharts radar chart with five axes.
Supabase: articles
Supabase stores the articles.
Visitor
Anyone comparing models.
Next.js app
A Next.js App Router app with shadcn/ui components.

Comparing two models

The versus view, step by step.

Pick two models1Look up both2Reasoning, knowledge, math, coding,creativity3Draw both on one radar4Two overlaid pentagons5What each model does better6VisitorCompare pageModel dataRadar chart
All 6 steps as text
What each part does
Visitor
A visitor.
Compare page
The compare page.
Model data
The model data.
Radar chart
The radar chart.

Decisions and trade-offs

  1. From a custom backend to a managed one

    It started as a monorepo with a FastAPI backend, ETL workers, Postgres and Redis in Docker, and was simplified to a Next.js app on Supabase: less to run, at the cost of the custom API.

  2. A versus view over tables

    A radar chart with two overlaid models was chosen over tabbed tables after several interface iterations, because the difference is visible at a glance.

  3. Say what is real

    The model data is demo data and the extraction worker is a prototype. This page says so, rather than implying a live pipeline.

Results

  • 5comparison axesreasoning, knowledge, math, coding, creativity
  • 26commits25 of them in two weeks of 2025
  • 1stproject of this setwhere the later work started

Stack

App
  • Next.js
  • TypeScript
  • Tailwind CSS
  • shadcn/ui
  • Recharts
Data
  • Supabase
  • AgentQL (prototype)