Skip to content

Available · Berlin · Werkstudent from Sept 2026

ArchitectingintelligentSystems.

I build AI systems that run in production — a phone receptionist you can call right now, an agent fleet that repairs itself overnight, and retrieval pipelines with the eval gates and cost tracing to prove they work. CS student in Berlin, open to Werkstudent roles.

01 — Proof

The difference between a demo and a system is measurability.

38%

Token cost per message, cut and measured

1,118 → 695 tokens · hybrid RAG replaced prompt-stuffing · traced in Langfuse

0.8%

Hard-failure rate across 520 agent runs

lease-queue dispatcher, 17 agents, self-healer · a 55% regression diagnosed and brought to 0%

12/12

Voice receptionist outcome eval, golden calls

offline rubric, 2026-08-25 · AI disclosure + consent evidence logged on every real call

Live demo

Don't take my word for it — hear it.

The receptionist answers in German or English, books a table or an appointment, and handles interruptions without losing the thread. It tells you it's an AI in the first sentence (EU AI Act Art. 50) and asks before anything is recorded (§201 StGB). Both are written to a tamper-evident journal on every call.

Live voice demo on request — email me and Clara will be on the line within the hour.

Demo on request · you're talking to an AI

02 — Selected Work

Selected Work

01

AI Phone Receptionist

Live · German & English · Vapi

Owner-run businesses lose bookings to phones nobody answers. The live demo answers for a Berlin estate agency: it qualifies buyers and renters in English or German, saves every caller's details to the database at hang-up, and books through a tool webhook. It discloses that it's an AI in its first sentence, asks consent before recording, and writes hash-chained evidence for each conversation — the parts a German business actually needs before it can use one.

Live demo — call it · 12/12 outcome eval (2026-09-02, offline) · compliance evidence per callVapi · Deepgram · ElevenLabs · Claude · FastAPI · PythonCall the demo
02

Self-Healing Agent Fleet

Self-built · Operations

Seventeen agents run the server from a Postgres-leased queue — code review, security sweeps, backup verification, evals, log digests. A classify → policy → remediate loop repairs failed runs overnight, a cost-truth ledger enforces a daily cap, and a morning Telegram digest reports what happened. When a model-routing change broke 55% of runs, the fleet's own logs led to the fix in one afternoon.

520 runs since Jul 9 · 0.8% hard failures · 55% → 0% regression fixPython · PostgreSQL · PM2 · Claude Code
03

Multilingual RAG Commerce Agent

Pilot project · Commerce · WhatsApp

A WhatsApp sales assistant over a furniture catalogue, built for a family business as a pilot. Hybrid retrieval — a keyword pass catches SKUs that embeddings dilute — replaced dumping the catalogue into every prompt. A semantic cache (95% cosine, 7-day TTL) absorbs repeat questions, and every response is cost-traced in Langfuse. Answers in English, Urdu and Roman Urdu.

38% token cost cut (1,118 → 695 / message) · Langfuse-traced · 10/10 retrieval evalPython · FastAPI · ChromaDB · Claude · Langfuse · Meta WhatsApp API
04

Autonomous Job Pipeline

Self-built · Automation · running daily

A five-stage pipeline that runs itself every morning: scrape Adzuna, Arbeitnow and Firecrawl, score each posting by weighted fit, tailor a CV and cover letter per match with Claude, and send a Telegram digest — degrading gracefully when a source is down. It found the role you might be hiring for.

Runs daily on cron · ~230 postings per run · tailored CV + letter per matchPython · Claude · Firecrawl · cronGitHub

Evidence · generated 2026-09-02

Eval gates, run counts, and one live pipeline run — dated.

12/12

Voice receptionist — phase-7 outcome QA (golden calls)

2026-09-02 · live

make eval → uv run python tests/eval_phase7.py (offline, deterministic rubric)

10/10

Restaurant bot — RAG retrieval (gold questions)

2026-09-02 · offline

make eval → uv run python tests/eval.py (offline, deterministic; chroma reindexed first)

10/10

Sales OS — missed-call exposure scorer

2026-09-02 · pipeline (manual runs)

make eval → uv run python tests/eval_exposure.py (offline fixtures)

0.8%

hard-failure rate across 520 agent runs

since 2026-07-09 · 17 agents enabled

343 done · 173 skipped · 4 failed

5/20

prospects from a live Sales OS run — Friseursalon in Berlin Neukölln

2026-08-25

20 businesses scored on missed-call exposure · ~$0.83 API cost · 0 cold messages sent (UWG §7 — consent first)

03 — How I Build

How I build systems that ship.

  1. Understand

    Define the real problem and what “working” means in numbers before writing code.

  2. Architect

    Design retrieval, agents, data flow, and failure modes together — not failure modes last.

  3. Build & evaluate

    Ship with an eval gate in CI, so quality is proven on every commit, not claimed once.

  4. Observe & iterate

    Trace every LLM call in production; tune on real usage and real cost.

04 — Stack

The tools I reach for.

Languages: Python, TypeScript, SQL, Bash
AI & retrieval: Claude, OpenRouter, ChromaDB, MiniLM embeddings, hybrid search, semantic caching, eval gates
Backend: FastAPI, Pydantic, Uvicorn, slowapi
Data & infra: PostgreSQL, Supabase, Redis, ClickHouse, Docker, Caddy, Cloudflare, PM2
Observability: Langfuse, structured logging
Voice & messaging: Vapi, Deepgram, ElevenLabs, WhatsApp Cloud API, Telegram Bot API
Automation: n8n, Playwright, Firecrawl
Frontend: Next.js, React, GSAP, Three.js
Quality: GitHub Actions, pytest, ruff, gitleaks

05 — About

The person behind the systems.

I'm Shehryar, a computer science student at Arden University in Berlin. For the past four months I've built and operated a production AI stack on my own server — Linux and Docker underneath, RAG and eval layers in the middle, and the frontend you're reading now on top.

I work eval-first. The part that separates a demo from a system is evaluation, tracing, and cost — so every project here ships with a dated eval result and a Langfuse trace behind it. When something breaks, I'd rather fix the cause than the symptom: when a bot token showed up in a log, I fixed the logger, not just the key.

When a product needs to speak more than one language — English, German, Urdu — I build that too.

Open to Werkstudent roles in Berlin. Remote-capable now, in Berlin full-time from 16 September.

Currently

Location
Berlin
Studying
BSc Computer Science, Arden University
Focus
AI systems · Full-stack
Open to
Werkstudent roles

06 — Contact

Let's build something that works.

Open to Werkstudent roles in Berlin — or just say hello.