ENGINEER
Description
I build multi-agent systems and LLM-powered products: Claude Agent SDK and MCP orchestration, retrieval pipelines, fine-tuning and evaluation, shipped on Cloud Run and Azure.
Hello, I’m
Forward Deployed
Forward Deployed AI Engineer, AI Architect, AI Designer
AI Engineer building multi-agent systems and LLM-powered products with the Claude Agent SDK and MCP, grounded in deep learning. Always up for AI research collaborations.
I build multi-agent systems and LLM-powered products: Claude Agent SDK and MCP orchestration, retrieval pipelines, fine-tuning and evaluation, shipped on Cloud Run and Azure.
I ship what I engineer — an iOS app on a live 3D globe, a Flutter app live on the App Store, and web products that stay up.
4,000+ hours in Claude Code. I run software like a tech lead whose team is AI agents: written contracts, isolated worktrees, adversarial review, and nothing called done without evidence.
Always evolving · September 2026
New models and agent tools arrive every few weeks. I test them as they land, keep what makes the work better, and change how I work — staying current is part of the job. What never changes: a written plan, independent review, and proof before anything ships.
Workflow changelog
A brief becomes a plan, split into packages that never touch the same file.
The main session is an orchestrator. It plans at full effort, and every package names the files it owns and the files it must leave alone.
The orchestrator writes the workflow. The graph is what runs.
A planner drafts, a critic attacks the draft, and the revised plan fans out to executors. Architecture calls get a read-only second opinion from another model family.
Executors work in parallel, each in its own git worktree, each under a written contract.
Every brief says what done means, when to stop and ask, and which command proves it. Executors cannot spawn agents, add dependencies, commit or push.
A reviewer re-runs every command and blocks what would ship wrong.
Findings come with a file, a line, why it is wrong and how to show it fails. At most two fix rounds, then the orchestrator decides.
Nothing is done until the orchestrator re-runs the checks. Every call lands in the ledger.
Agent reports are claims. The acceptance commands are run again before any merge, and each non-obvious decision becomes a dated ledger row with its reason and its undo.

The tools I build with, and how I use them: Claude Code plans and edits, Codex gives a read-only second opinion, and every technology below is tied to the projects that run it.

Logos are trademarks of their owners.
Depth of experience = projects × time. % CPU here means depth of experience, not a measurement: 70% from how many of my 13 projects and roles use it (against the busiest), 30% from how long I have used it (against the longest). Counted from my case studies and work history, September 2026.
Also installed, not sampled: PyTorch, TensorFlow.
Illustrative session, written for this site.
Claude Code adds a citation guard to the chat API, so no sentence ships without a passage pgvector retrieved, and the tests pass. Codex, read-only, flags that one uncited summary sentence would now fail the whole answer. Claude Code turns that case into a warning and the tests pass again.
Claude Code > No sentence in an answer ships without a retrieved source. ⏺ Update(app/citations.py) 24 + if any(not markers(s) for s in sentences(answer)): raise Uncited() ⎿ pytest -q: 7 passed Codex (read-only second opinion) › Second opinion on the citation guard. Read-only. • Explored citations.py, chat.py, test_citations.py • citations.py:24 fails the whole answer on one uncited summary sentence. Models write those; warn, don't block. ─ Worked for 11s ─ Claude Code > Codex: warn on uncited sentences, don't block the answer. ⏺ Update(app/citations.py) 24 + warnings += [s for s in sentences(answer) if not markers(s)] ⎿ pytest -q: 8 passed
Claude Code animates the tier-up banner on the run summary in SwiftUI and runs the tests. Codex, read-only, flags that the spring still runs when Reduce Motion is on. Claude Code switches to a short fade in that case and the tests pass.
Claude Code > Animate the tier-up banner on the run summary. ⏺ Update(TierUpBanner.swift) 20 + .animation(.spring(duration: 0.4), value: tier) ⎿ xcodebuild test: 14 passed Codex (read-only second opinion) › Review the banner change. Read-only. • Explored TierUpBanner.swift • TierUpBanner.swift:20 still springs with Reduce Motion on. Read accessibilityReduceMotion and fade instead. ─ Worked for 8s ─ Claude Code > Codex: honour Reduce Motion. ⏺ Update(TierUpBanner.swift) 20 + .animation(reduce ? .easeOut : .spring(duration: 0.4), value: tier) ⎿ xcodebuild test: 15 passed
The director session in Claude Code turns the picked proposal into a brief for tomorrow's carousel. Codex renders a still for the first slide and flags that the brief asks for a drawn app screen. The director replaces it with a request for a real capture; nothing is published without approval.
Claude Code > Brief tomorrow's carousel from the proposal I picked. ⏺ Write(jobs/next-carousel/brief.md) ⎿ Wrote 38 lines, waiting for your approval Codex (read-only second opinion) › Still for slide 1: a runner at first light, 4:5, no text. • Generated still-01.png, 1080 × 1350 • brief.md:14 draws an app screen on slide 3. Only real app UI: ask for a capture instead. ─ Worked for 34s ─ Claude Code > Codex is right. Slide 3 uses a real capture. ⏺ Update(jobs/next-carousel/brief.md) ⎿ Slide 3 is now a capture request for the app-capture agent
Claude Code swaps the desk's material in the three.js hero when the camera dives into the screen. Codex, read-only, flags that the old material and its texture are never disposed, so every scroll back leaks GPU memory. Claude Code disposes them and a test pins the texture count.
Claude Code > Swap the desk material when the camera dives in. ⏺ Update(src/hero/desk.ts) 43 + desk.material = diveMaterial(walnut, DIVE_TEXTURE) ⎿ vitest: 6 passed Codex (read-only second opinion) › Second opinion on this diff. Read-only. • Explored desk.ts, heroController.ts • desk.ts:43 replaces the material but never disposes the old one or its texture. Each scroll back leaks GPU memory. ─ Worked for 14s ─ Claude Code > Codex: dispose what you replace. ⏺ Update(src/hero/desk.ts) 42 + desk.material.map?.dispose(); desk.material.dispose() ⎿ vitest: 7 passed, texture count stays at 12 after 20 dives
Claude Code makes the discovery loop crawl again with model-proposed queries until a judge finds each university's admission call. Codex, read-only, flags that one failed crawl would end the whole multi-university run. Claude Code logs and skips failures and the run completes.
Claude Code > Keep crawling until the judge finds each admission call. ⏺ Update(discovery/loop.py) 32 + found += crawl(model.propose_queries(uni, found)) ⎿ pytest -q: 9 passed Codex (read-only second opinion) › Read-only review of the discovery loop. • Explored loop.py, test_discovery.py • loop.py:32 lets one failed crawl raise and end the whole multi-university run. Log it and keep going. ─ Worked for 10s ─ Claude Code > Codex: one flaky call must not stop the run. ⏺ Update(discovery/loop.py) 32 + found += crawl(model.propose_queries(uni, found), on_error=skip) ⎿ pytest -q: 10 passed, still-missing universities are flagged
Claude Code gives the copywriter agent a fact-check tool and registers it. Codex, read-only, flags that the agent's identity file still lists only its old tools, so it may never call the new one. With the owner's approval the tool is added there, and a live, instrumented run proves the call happens.
Claude Code > Give the copywriter a tool to check claims against the research. ⏺ Update(sdk/tools/research.py) ⎿ check_claim registered on the copywriter; 12 tests passed Codex (read-only second opinion) › Is the new tool really wired? Read-only. • Explored agents.py, copywriter/identity.md • identity.md:9 still names only the old tools. Registered and tested, but the agent may never call it. ─ Worked for 12s ─ Claude Code > Approved: add it to the identity file, then prove the call. ⏺ Bash(python probe.py --agent copywriter --watch tools) ⎿ check_claim called 3 times in one live run
Claude Code reads a Figma frame through the Figma MCP server and writes the React component. Codex, read-only, flags hard-coded colours and radii where the frame uses design tokens. Claude Code switches to the tokens and the tests pass.
Claude Code > Build the pricing card from the selected Figma frame. ⏺ figma · get_design_context (MCP) ⎿ Wrote PricingCard.tsx from 3 variants and 11 tokens Codex (read-only second opinion) › Second opinion on PricingCard.tsx. Read-only. • Explored PricingCard.tsx, tokens.css • PricingCard.tsx:22 hard-codes #0B1015 and 14px where the frame uses tokens. Take them from the design context. ─ Worked for 9s ─ Claude Code > Codex: tokens, not literals. ⏺ Update(src/components/PricingCard.tsx) 22 + background: "var(--surface-1)", borderRadius: "var(--radius-lg)", ⎿ npm test: 5 passed
Claude Code adds a second Gemini model that re-solves exam questions with shuffled options and flags a question when the two disagree. Codex, read-only, flags that it runs on every question, when only the blind solver's low-confidence ones need it. Claude Code gates it on confidence.
Claude Code > Add a second model that re-solves the uncertain questions. ⏺ Update(validation/cross_model.py) 12 + if second_model.solve(item.shuffled()) != item.key: item.flag() ⎿ pytest -q: 23 passed Codex (read-only second opinion) › Review the cross-model check. Read-only. • Explored cross_model.py, validator.py • cross_model.py:12 re-solves every question. Only the blind solver's low-confidence ones need a second model. ─ Worked for 13s ─ Claude Code > Codex: gate it on confidence. ⏺ Update(validation/cross_model.py) 11 + if item.confidence >= settings.cross_model_below: return ⎿ pytest -q: 24 passed, the second model runs on low confidence only
Claude Code writes the monthly job that re-scrapes the course pages and re-embeds only what changed, reusing the service's container image. Codex, read-only, flags that a site's old embeddings are deleted before its new scrape succeeds. Claude Code replaces them only after a successful scrape.
Claude Code > Rebuild the course corpus every month without breaking it. ⏺ Update(jobs/refresh.ts) 57 + await store.clear(site); await store.upsert(site, chunks) ⎿ npm test: 11 passed Codex (read-only second opinion) › Read-only review of the refresh job. • Explored refresh.ts, store.ts • refresh.ts:57 clears a site before knowing its scrape worked. A failed site ends up with an empty corpus. ─ Worked for 12s ─ Claude Code > Codex: failure may only keep things, never delete them. ⏺ Update(jobs/refresh.ts) 57 + if (scraped.ok) await store.replace(site, chunks) ⎿ npm test: 12 passed, a failed site keeps last month's corpus
Claude Code makes the home screen's BLoC fetch today's prayer times when it opens, and the Flutter tests pass. Codex, read-only, flags that it asks again on every open, even when the day and the area have not changed. Claude Code keeps the saved times and fetches only on a new day or a new place.
Claude Code > Load today's prayer times when the home screen opens. ⏺ Update(lib/home_bloc.dart) 41 + emit(HomeLoaded(await repo.fetch(today, area))) ⎿ flutter test: 12 passed Codex (read-only second opinion) › Review the home screen change. Read-only. • Explored home_bloc.dart, prayer_repository.dart • home_bloc.dart:41 asks again on every open, even when the day and the area are unchanged. Reuse the saved times. ─ Worked for 10s ─ Claude Code > Codex: fetch only on a new day or a new place. ⏺ Update(lib/home_bloc.dart) 41 + if (!saved.matches(today, area)) saved = await repo.fetch(today, area) ⎿ flutter test: 13 passed, a reopen on the same day asks nothing
Claude Code prepares an instruction set for a fine-tune: augment, then split into train and eval. Codex, read-only, flags that splitting after augmentation puts paraphrases of one item on both sides, so the evals look better than they are. Claude Code splits first and augments the training side only.
Claude Code > Prepare the instruction set for the next fine-tune. ⏺ Update(pipeline/prepare.py) 18 + train, evals = split(augment(rows), 0.95, seed=7) ⎿ pytest -q: 9 passed Codex (read-only second opinion) › Anything wrong with the split? Read-only. • Explored prepare.py • prepare.py:18 augments before the split, so paraphrases of one item land in train and eval. Evals will look too good. ─ Worked for 9s ─ Claude Code > Codex: split first, augment the training side only. ⏺ Update(pipeline/prepare.py) 18 + train, evals = split(rows, 0.95, seed=7); train = augment(train) ⎿ pytest -q: 10 passed, no eval item has a twin in train
Claude Code is a product of Anthropic. Codex is a product of OpenAI. Used to show the tools I work with.
Territory running app for iPhone, live on the App Store: close a loop and the streets inside become yours.
Ten specialised agents research, write, illustrate and review content end to end, across 50+ agent tools.
A Qur'an and prayer companion for iPhone and Android, co-built in a team of two and live on both stores.
Specialised agents turn source PDFs into exam questions, and a cross-model validator checks every item.
Italian-language RAG assistant, deployed to the University of Ferrara course pages, answering with cited sources.
An agentic crawl extracts admission rules with evidence checks, feeding a chatbot that cites its sources.
Compares LLMs head to head on a five-axis radar chart, with an AgentQL extraction prototype for new data.
Seven AI agent sessions run RouteRush's marketing like an agency, and a human approves every brief and post.
RouteRush's marketing site: a live 3D map you fly by scrolling, in six languages built from one English source.
Neurolanche X Labs
2024
Built fine-tuning and end-to-end pipelines, served models on Azure AI and AWS, designed the Flutter app’s backend architecture, and used STT/TTS for live voice chat.
Outlier
2025
Fine-tuned LLMs alongside AI researchers, ran systematic reasoning and grading evaluations, and built the augmentation and quality-control pipelines behind them.
SmartCreative SRL
2026
Moved the design-to-code workflow onto Figma MCP end to end, cutting frontend delivery from days to hours, and built the agentic workflows around the rest of the process.
Available
NOW
Put me on your team and I’ll ship the AI your customers actually need.Hire me