Skip to content
View technikky's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report technikky

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
technikky/README.md
github contribution grid snake animation

generated with technikky/snk

Nur Afzan Zulaikha

AI Engineer · LLM Evaluation & Benchmark Design · Full-Stack (Python / TypeScript)

I build and evaluate LLM systems: retrieval pipelines, rubric-based evaluation harnesses, and graded benchmark tasks for AI coding agents. I have authored 120+ reviewed AI-training tasks, including software-engineering benchmarks with hidden test suites in Python/pytest and TypeScript/vitest, calibrated against frontier models.

Based in Penang, Malaysia. Open to remote AI engineering and AI evaluation work.

What I work on

  • LLM evaluation — rubric design, deterministic scoring, judge calibration, regression suites
  • Retrieval — chunking, embeddings, vector search, reranking, and measuring whether it actually works
  • Benchmark design — unambiguous specifications, fail-to-pass / pass-to-pass tests, difficulty calibration
  • Full-stack delivery — FastAPI and Node services, Next.js front ends, Docker, GitHub Actions

Stack

Languages Python · TypeScript · JavaScript · SQL Backend FastAPI · Flask · Node.js · REST · GraphQL · WebSockets AI/ML PyTorch · TensorFlow · Hugging Face · LangChain · LoRA/QLoRA Retrieval Chroma · FAISS · Pinecone · Redis Frontend React · Next.js · Tailwind CSS · Storybook Infrastructure Docker · Kubernetes · AWS · GitHub Actions

Featured

Project What it does
llm-evaluation-framework Reproducible rubric-based LLM evaluation harness — deterministic scorers, LLM-judge with abstention, bootstrap confidence intervals
workforce-ops-dashboard Workforce operations dashboard — React, TypeScript, role-based views
offline-english-learning Offline-first AI English learning system for schools — Electron + Flutter

Contact

afzanzulaikha97@gmail.com

This GitHub account is recent — earlier work was in private employer repositories.

Pinned Loading

  1. ai-coding-agent-benchmarks ai-coding-agent-benchmarks Public

    Validated benchmark harness for AI coding agents - hidden graded tests, fail-to-pass/pass-to-pass validation, difficulty calibration. Python/pytest + TypeScript/vitest, Docker, GitHub Actions. All …

    Python

  2. ai-content-analyzer ai-content-analyzer Public

    Full-stack content analyzer - summarisation, sentiment, entity extraction, document Q&A, each with a measured baseline. Next.js 16 + FastAPI, deterministic JSON repair, dev/test split evaluation.

    Python

  3. fullstack-typescript-platform fullstack-typescript-platform Public

    Auth, RBAC, REST + GraphQL and WebSockets over one domain layer in TypeScript. Authorization enforced once and proven equal across both surfaces; PostgreSQL 17 in-process so 550 tests run real SQL …

    TypeScript

  4. llm-evaluation-framework llm-evaluation-framework Public

    Rubric-based LLM evaluation harness - an unparseable judge response abstains instead of scoring zero, every mean carries a seeded bootstrap interval, and rubrics are hashed into a config_hash. Pyth…

    Python

  5. llm-finetuning-lab llm-finetuning-lab Public

    LoRA/QLoRA fine-tuning with exact parameter and memory budgets cross-checked against transformers and peft, three tested adapter invariants, and leakage detection that refuses to train.

    Python

  6. production-rag-system production-rag-system Public

    RAG pipeline with measured retrieval quality - hybrid dense+BM25 with reciprocal-rank fusion, reranking, MMR, nDCG/MRR/Recall evaluation, citation grounding. FastAPI, Docker, GitHub Actions.

    Python