generated with technikky/snk
AI Engineer · LLM Evaluation & Benchmark Design · Full-Stack (Python / TypeScript)
I build and evaluate LLM systems: retrieval pipelines, rubric-based evaluation harnesses, and graded benchmark tasks for AI coding agents. I have authored 120+ reviewed AI-training tasks, including software-engineering benchmarks with hidden test suites in Python/pytest and TypeScript/vitest, calibrated against frontier models.
Based in Penang, Malaysia. Open to remote AI engineering and AI evaluation work.
- LLM evaluation — rubric design, deterministic scoring, judge calibration, regression suites
- Retrieval — chunking, embeddings, vector search, reranking, and measuring whether it actually works
- Benchmark design — unambiguous specifications, fail-to-pass / pass-to-pass tests, difficulty calibration
- Full-stack delivery — FastAPI and Node services, Next.js front ends, Docker, GitHub Actions
Languages Python · TypeScript · JavaScript · SQL Backend FastAPI · Flask · Node.js · REST · GraphQL · WebSockets AI/ML PyTorch · TensorFlow · Hugging Face · LangChain · LoRA/QLoRA Retrieval Chroma · FAISS · Pinecone · Redis Frontend React · Next.js · Tailwind CSS · Storybook Infrastructure Docker · Kubernetes · AWS · GitHub Actions
| Project | What it does |
|---|---|
| llm-evaluation-framework | Reproducible rubric-based LLM evaluation harness — deterministic scorers, LLM-judge with abstention, bootstrap confidence intervals |
| workforce-ops-dashboard | Workforce operations dashboard — React, TypeScript, role-based views |
| offline-english-learning | Offline-first AI English learning system for schools — Electron + Flutter |
This GitHub account is recent — earlier work was in private employer repositories.



